scanmydocz Nothing leaves this browser

Tools / PDF OCR

PDF OCR

Take the words out of a PDF as a plain text file.

  • No upload endpoint exists
  • No network code in any engine
  • Every tool driven in a browser each release

Choose a PDF

or drop them here

How it works

  1. Add the PDF.
  2. Choose whether to keep the line breaks and whether to mark page starts.
  3. Take the text out, then download the .txt file.

When the browser is not enough

scanmydocz for Android

Live edge detection through the camera, text recognition that makes the PDF searchable, a signature you can place, and batch capture. All of it on the phone.

In closed testing. There is no public Play link yet, so there is no button here yet.

scanmydocz Pro

Pro is the paid tier in the Android app. It cannot be bought on the web yet, because checkout is not connected, and we would rather say that than show you a price you cannot pay.

Questions

It says there is no text in my PDF. Why?

Because its pages are pictures. That is what a scan or a photographed document is, and a picture of words contains no words, however clearly you can read them. Getting text out of that needs recognition rather than extraction, which is a different job. The page tells you this before it hands you anything, rather than giving you an empty file to puzzle over.

Do my files get uploaded?

No. The PDF is read in this page, in your browser. There is no upload step and no server here that could receive one.

Will the text come out in the right order?

Usually, and it is not guaranteed. A PDF stores letters at positions and has no idea what a column is, so a two column layout or a newspaper page can interleave. Ordinary documents come out correctly; check anything that was laid out in columns.

What happens to tables?

They lose their shape. A table in a PDF is text at coordinates with no structure behind it, so it arrives as lines of words. There is no honest way around that at this level.

What is the difference between the two layout choices?

Keeping line breaks gives you the text as the page laid it out, one line per line. Joining paragraphs removes the line wraps inside a paragraph but keeps the breaks between them, which is what you want for pasting into something that will re-wrap the text itself.

Can I get the text of one page only?

Not directly. Take that page out with Split PDF first, then run it through here.

Does it work without a connection?

Opening the page needs one, and so does the first PDF you open, because that is when this site hands over the reader. Everything after that is local.

Other tools