Tools / PDF OCR
PDF OCR
Take the words out of a PDF as a plain text file.
- No upload endpoint exists
- No network code in any engine
- Every tool driven in a browser each release
Choose a PDF
or drop them here
How it works
- Add the PDF.
- Choose whether to keep the line breaks and whether to mark page starts.
- Take the text out, then download the .txt file.
When the browser is not enough
scanmydocz for Android
Live edge detection through the camera, text recognition that makes the PDF searchable, a signature you can place, and batch capture. All of it on the phone.
In closed testing. There is no public Play link yet, so there is no button here yet.
scanmydocz Pro
Pro is the paid tier in the Android app. It cannot be bought on the web yet, because checkout is not connected, and we would rather say that than show you a price you cannot pay.
Questions
It says there is no text in my PDF. Why?
Because its pages are pictures. That is what a scan or a photographed document is, and a picture of words contains no words, however clearly you can read them. Getting text out of that needs recognition rather than extraction, which is a different job. The page tells you this before it hands you anything, rather than giving you an empty file to puzzle over.
Do my files get uploaded?
No. The PDF is read in this page, in your browser. There is no upload step and no server here that could receive one.
Will the text come out in the right order?
Usually, and it is not guaranteed. A PDF stores letters at positions and has no idea what a column is, so a two column layout or a newspaper page can interleave. Ordinary documents come out correctly; check anything that was laid out in columns.
What happens to tables?
They lose their shape. A table in a PDF is text at coordinates with no structure behind it, so it arrives as lines of words. There is no honest way around that at this level.
What is the difference between the two layout choices?
Keeping line breaks gives you the text as the page laid it out, one line per line. Joining paragraphs removes the line wraps inside a paragraph but keeps the breaks between them, which is what you want for pasting into something that will re-wrap the text itself.
Can I get the text of one page only?
Not directly. Take that page out with Split PDF first, then run it through here.
Does it work without a connection?
Opening the page needs one, and so does the first PDF you open, because that is when this site hands over the reader. Everything after that is local.
Other tools
-
PDF Scanner
Turn a photo of a document into a straight, clean PDF.
-
JPG to PDF
Put images into a PDF, in the order you choose.
-
PDF to JPG
Get every page of a PDF as an image.
-
Compress PDF
Make a PDF smaller, and see exactly what it cost.
-
ID Card Scan
Front and back of a card, on a single A4 sheet.
-
Merge PDF
Join several PDFs into one, in the order you choose.
-
Split PDF
Take a range of pages out, or break a file into its pages.
-
Rotate PDF
Turn the pages the right way up, without redrawing them.
-
Organise PDF
See every page, drop the ones you do not want, move the rest.
-
Protect PDF
Put a password on a PDF, with AES 256, in your browser.
-
Unlock PDF
Take the password off a PDF you can already open.
-
Edit PDF
Add text, cover something, drop a signature on it.