Making a scanned PDF searchable

OCR (optical character recognition) reads the pixels of a scanned page and produces a text layer, so the document becomes searchable and copyable. The recognition runs in your browser.

How to do it

  1. Add the scanned PDF. Choose the file. Recognition quality depends on scan quality: 300 dpi or better, straight, and evenly lit gives the best result.
  2. Choose the language. Select the language of the document. Choosing the wrong one is the most common cause of poor output, because the recogniser looks for the wrong character set.
  3. Run recognition. The engine processes each page and writes a text layer behind the image. The page still looks identical; the text is now selectable.
  4. Export and check. Save the searchable PDF. Open it and try selecting a line of text to confirm the layer is present.

What OCR can and cannot fix

OCR converts an image of text into actual text. It does not improve a bad scan. A skewed, blurred or low-resolution page produces a text layer with errors, and a document photographed with a phone at an angle will recognise worse than the same page scanned flat.

Accuracy also depends on the material. Clean printed text in a common language recognises well. Handwriting, ornate fonts, tables with thin rules, and multi-column layouts are substantially harder and should be checked before being relied on.

OCR before conversion, not after

If your goal is an editable Word file, the order matters. A scanned PDF has no text to convert, so converting it directly produces an empty or image-only document. Run OCR first to create the text layer, then convert. That two-step route is the only one that works on scanned material.

Privacy in recognition

OCR is often the most sensitive operation a person performs, because scanned documents tend to be identity papers, medical records and signed agreements. Here the recognition model runs locally, so the page images are not transmitted for processing.

What this tool does not do

Stated plainly, because a limitation you discover after trusting the result is worse than one you read first.

Frequently asked questions

How do I convert a scanned PDF to searchable text?

Run OCR on it. The recogniser reads each page image and writes a text layer behind it, after which the document can be searched and copied like any other PDF. Convert to Word afterwards if you need an editable file.

Can I convert a scanned PDF to Word for free?

Yes, but in two steps: OCR first to create the text layer, then convert the searchable PDF to Word. Converting a scanned PDF directly will not work, because there is no text in it to convert.

What resolution do I need for good OCR?

300 dpi is the usual recommendation for printed text. Below roughly 200 dpi, recognition accuracy falls sharply. Higher than 600 dpi rarely helps and makes processing slower.

Does OCR work on handwriting?

Poorly. Handwriting recognition is a different and much harder problem than printed-text recognition. Expect to correct the output by hand.

Can I OCR a PDF on my phone?

Yes, and it is often convenient, since the phone camera is the scanner. For best results hold the camera square to the page and use even lighting rather than a flash.