How to convert a scanned PDF to text (OCR)

A scanned PDF looks like a document, but to the computer it is a stack of pictures: you cannot select, search or copy anything. OCR reads the letters in those images and turns them into real text. What it takes for it to get things right, and how to do it without uploading the document to any website.

Reviewed on 3 October 2026 · 4 min read

First, check whether your PDF needs OCR

Not every PDF is an image. To find out, open the file and try to select a word with the mouse, or press Ctrl+F (Cmd+F on a Mac) and search for a word you know is in the document. If the text selects and the search finds it, the PDF already has a text layer and you do not need OCR: just extract it with the PDF text extractor. If nothing can be selected, it is an image and OCR is the way.

What OCR does

OCR (optical character recognition) analyses the image of each page, identifies the shape of the letters and turns them into text. It is the job of phone scanner apps. UtilsDock's PDF OCR tool draws each page as an image at a resolution good enough for small print, passes it through a recognition engine (Tesseract.js) that runs on your own device, and joins the text from all pages into one block you can copy or download. The images from the PDF are not uploaded to any server, which matters if it is a contract, a payslip or a document with personal data.

The first time takes a little longer, because the engine and the language data (a few megabytes) download and the browser keeps them; after that it is quicker. If what you have is a photo rather than a PDF, the image to text tool uses the same engine.

How to get it right: what matters most

OCR accuracy depends mostly on the quality of the starting image. The documentation for Tesseract, the most widely used open-source OCR engine, lists the factors that weigh most:

  • Resolution. Tesseract works best on images of at least 300 dpi, and recommends enlarging smaller ones. A 200 dpi scan can do if the type is large and sharp, but for small print or thin fonts, 300 dpi gives better results.
  • Straight text. A page scanned crooked significantly reduces the quality of line segmentation, and with it the quality of the OCR. If your scan came out tilted, straighten it first with the scanned document straightener.
  • Dark text on a light background. This is what the engine reads best; coloured backgrounds or pages shadowed by a book spine make things worse.
  • No dark borders. The black borders scanners leave around the page can be picked up by mistake as characters. Crop the image to the text area with a sensible margin.
  • No noise. Noise (random variation in brightness or colour) makes the text harder to read. A clean, well-lit scan with no creases reads far better than a photo taken at an angle in poor light.

Language matters

The engine uses a different model for each language. The tool applies the one for the language of the page: on this version, English. A document in another language reads worse, with accents and special characters being the first casualties. If your document is in a language other than the page's, expect more errors and check more carefully.

What to expect from the result

  • Clean printed type: very good, with the odd isolated error.
  • Handwriting: far less reliable. OCR is built for printed text.
  • Tables and columns: the text may come out in a strange order or with columns mixed together.
  • Numbers and codes: a 0 and an O, or a 1 and an l, are easily confused. Check any amount, date, document number or IBAN against the original.

So do not rely on the result for anything important without reviewing it. A quick pass with the original beside you catches nearly every error.

What this tool does not do

It gives you the text, not a PDF with the original images and an invisible text layer added back, which is the format desktop programs produce and what makes a PDF searchable. That is a more complex file to build than a browser can reliably produce. If you need a "searchable" PDF for archiving, use a PDF program that does it; if you need the content to copy, quote or edit, the text is enough.

Before you start: prepare the document

  1. Scan at 300 dpi if you can, or rescan only the pages that come out badly.
  2. Check that pages are straight, with good contrast and no shadows.
  3. If the PDF is very heavy, do not compress it first: aggressive compression erases detail and OCR notices.
  4. When it finishes, compare the text with the original on the critical parts.

If the document holds sensitive data, remember the whole process runs in your browser; after OCR you can also redact the details you do not want to share before sending it.

Sources and further reading

Figures checked on 3 October 2026.

Do it now, free, in your browser. Your files are not uploaded.

Read the text in a scanned PDF and get it as plain, searchable text.

Frequently asked questions

How do I know if a PDF is an image?
Try to select a word with the mouse or search for one with Ctrl+F. If nothing can be selected or found, the PDF is an image and needs OCR.
How do I convert a scanned PDF to text for free?
With an OCR tool that runs in the browser: open the PDF, wait for it to recognise the pages, then copy or download the text. With UtilsDock the PDF never leaves your device.
What resolution does OCR need to be accurate?
Tesseract's documentation says it works best on images of at least 300 dpi. With large, sharp type, 200 dpi may be enough.
Why does OCR confuse letters and numbers?
Because some shapes are very similar (0 and O, 1 and l) and because scan quality has a big effect. Improve resolution, contrast and straightness of the page, and always check figures and codes.
Does OCR work on handwriting?
Much worse than on printed text. The engine is built for printed type; with handwriting, expect many errors.
Do I get a searchable PDF?
No: the tool returns text, not a PDF with an invisible text layer. To archive a searchable PDF you need a desktop PDF program.