How to use OCR PDF
- Choose a scanned PDF.
- Pick the Language of the text. Choose English + Hindi or English + Telugu for forms that mix scripts.
- Leave Pages as all, or enter a range such as 1-3.
- Keep Scan resolution at 300 DPI for accuracy, or choose 200 DPI to finish faster.
- Press Make searchable. When it finishes, download the copy ending in -searchable.pdf, and copy or save the recognised text if you need it.
What it does and when to use it
A scanned PDF is a stack of pictures. You can read it, but your computer cannot: searching for a word finds nothing, you cannot copy a sentence, and screen readers have nothing to read aloud. OCR (optical character recognition) fixes that by recognising the letters in each picture.
This tool keeps your pages exactly as they look and lays the recognised words invisibly on top, each one positioned over the word in the image. In any PDF reader you can then press Ctrl+F (or Cmd+F) to search, drag to select a paragraph, or copy a reference number into an email.
Common uses include old letters and records scanned for an archive, a scanned contract you need to quote from, receipts and bills you want to search later, textbook pages for revision notes, and bilingual government forms with English and Hindi or Telugu side by side.
How it works
The tool uses tesseract.js 7, a WebAssembly version of the open-source Tesseract OCR engine. The first time you run it, the engine and the language files you chose download from jsDelivr, a public code host, and your browser caches them. Your PDF is never part of that download.
For each page, pdf.js renders the page as an image at the scan resolution you picked. Tesseract finds lines and words and returns each word with its position and a confidence score. pdf-lib then writes those words onto the original page as invisible text, using a special font with no visible shapes and a Unicode map. That is the same technique Tesseract’s own PDF output uses, and it is why Hindi and Telugu search works as well as English.
When it is done you get three things: the searchable PDF, the number of words recognised with their average confidence, and the plain recognised text page by page, which you can copy or download as a .txt file.
Worked examples
A three-page scanned letter in English. Choose English, leave pages as all and keep 300 DPI. The first run downloads about 6.9 MB, and then each page is recognised in turn with a progress message. The result looks identical to the scan, but searching for a name in your PDF reader now jumps to it.
A bilingual form in English and Hindi. Choose English + Hindi so both scripts are recognised in one pass. Mixed-script pages are harder, so check names, numbers and dates against the original.
A 40-page scanned book chapter. Consider 200 DPI to save time, or run a few pages first with a range such as 1-5 to see whether the quality is good enough before doing the whole chapter.
Limits and tips
- OCR makes mistakes, especially with numbers, small print, unusual fonts, stamps and handwriting. Proofread anything important, such as amounts or ID numbers.
- Straight, well-lit scans give much better results. Fix sideways pages with the PDF Rotator before running OCR.
- If you only need the words and not a PDF, the recognised text panel gives you plain text. For PDFs that already have text, PDF to Text is faster.
- To turn the searchable result into an editable document, run it through PDF to Word.
- If the searchable PDF is over an upload limit, compress it with lossless settings first. Compress PDF to 1 MB with image conversion off keeps the text layer; converting pages to images would remove it.
- Recognition is heavy work. Long documents run faster on a laptop or desktop than on a phone.
Frequently asked questions
Does the OCR engine send my PDF anywhere?
How big is the first download?
Can it read handwriting?
Why does my PDF already have selectable text?
Will the file get bigger?
What does the average confidence number mean?
How do we know nothing is uploaded? Test it yourself on the privacy proof page.