OCR PDF — make a scanned PDF searchable (English, Hindi, Telugu)

OCR PDF reads the words in a scanned PDF and adds an invisible text layer, so you can search, select and copy them while the pages look exactly the same.

Live demo · real result from this tool

Drop your PDF here

or paste with Ctrl+V · nothing is uploaded · add several to process them in one go

The OCR engine and language data download once from jsDelivr and are then cached by your browser. Your PDF is not uploaded.
For example 1-3, 5, 8- (8 to the end), odd, even or all.

Your result appears here — processed on your device.

Processed on your device — nothing is uploaded.

How to use OCR PDF

  1. Choose a scanned PDF.
  2. Pick the Language of the text. Choose English + Hindi or English + Telugu for forms that mix scripts.
  3. Leave Pages as all, or enter a range such as 1-3.
  4. Keep Scan resolution at 300 DPI for accuracy, or choose 200 DPI to finish faster.
  5. Press Make searchable. When it finishes, download the copy ending in -searchable.pdf, and copy or save the recognised text if you need it.

What it does and when to use it

A scanned PDF is a stack of pictures. You can read it, but your computer cannot: searching for a word finds nothing, you cannot copy a sentence, and screen readers have nothing to read aloud. OCR (optical character recognition) fixes that by recognising the letters in each picture.

This tool keeps your pages exactly as they look and lays the recognised words invisibly on top, each one positioned over the word in the image. In any PDF reader you can then press Ctrl+F (or Cmd+F) to search, drag to select a paragraph, or copy a reference number into an email.

Common uses include old letters and records scanned for an archive, a scanned contract you need to quote from, receipts and bills you want to search later, textbook pages for revision notes, and bilingual government forms with English and Hindi or Telugu side by side.

How it works

The tool uses tesseract.js 7, a WebAssembly version of the open-source Tesseract OCR engine. The first time you run it, the engine and the language files you chose download from jsDelivr, a public code host, and your browser caches them. Your PDF is never part of that download.

For each page, pdf.js renders the page as an image at the scan resolution you picked. Tesseract finds lines and words and returns each word with its position and a confidence score. pdf-lib then writes those words onto the original page as invisible text, using a special font with no visible shapes and a Unicode map. That is the same technique Tesseract’s own PDF output uses, and it is why Hindi and Telugu search works as well as English.

When it is done you get three things: the searchable PDF, the number of words recognised with their average confidence, and the plain recognised text page by page, which you can copy or download as a .txt file.

Worked examples

A three-page scanned letter in English. Choose English, leave pages as all and keep 300 DPI. The first run downloads about 6.9 MB, and then each page is recognised in turn with a progress message. The result looks identical to the scan, but searching for a name in your PDF reader now jumps to it.

A bilingual form in English and Hindi. Choose English + Hindi so both scripts are recognised in one pass. Mixed-script pages are harder, so check names, numbers and dates against the original.

A 40-page scanned book chapter. Consider 200 DPI to save time, or run a few pages first with a range such as 1-5 to see whether the quality is good enough before doing the whole chapter.

Limits and tips

  • OCR makes mistakes, especially with numbers, small print, unusual fonts, stamps and handwriting. Proofread anything important, such as amounts or ID numbers.
  • Straight, well-lit scans give much better results. Fix sideways pages with the PDF Rotator before running OCR.
  • If you only need the words and not a PDF, the recognised text panel gives you plain text. For PDFs that already have text, PDF to Text is faster.
  • To turn the searchable result into an editable document, run it through PDF to Word.
  • If the searchable PDF is over an upload limit, compress it with lossless settings first. Compress PDF to 1 MB with image conversion off keeps the text layer; converting pages to images would remove it.
  • Recognition is heavy work. Long documents run faster on a laptop or desktop than on a phone.

Frequently asked questions

Does the OCR engine send my PDF anywhere?
No. Your browser downloads the recognition engine and language data from jsDelivr the first time, then runs it on your device. The PDF itself never leaves your computer or phone.
How big is the first download?
About 6.9 MB for English, 5.3 MB for Hindi, 5.6 MB for Telugu, 8.2 MB for English plus Hindi and 8.6 MB for English plus Telugu. Your browser caches it, so later runs start faster.
Can it read handwriting?
Not reliably. The engine is trained on printed text. Neat block capitals sometimes work, but joined handwriting usually comes out as errors.
Why does my PDF already have selectable text?
It was probably created on a computer rather than scanned, so it already has a text layer and does not need OCR. Running OCR on it adds a second, invisible layer that is usually worse than the original.
Will the file get bigger?
A little. The invisible text adds some data to each page, but the page images are kept exactly as they were.
What does the average confidence number mean?
It is the engine's own estimate of how sure it was about each word, averaged across the document. A low figure suggests a blurry or skewed scan and more mistakes to check.

Written by the ToolsRift team · Last updated

Pages are drafted with AI assistance and checked by automated tests. How we write and check pages.

Was this tool useful?

Report a problem

Start typing to search every tool.

↑ ↓ to moveEnter to openEsc to close