Make a scanned PDF searchable
Make a scanned PDF searchable, on your own device.
Or drop one anywhere on this page
What to know about OCR PDF
A scanned document is a photograph of some pages: you can see the words, and your computer cannot. Text recognition reads the picture and writes what it finds into the file as an invisible layer over each word, so the document becomes searchable while looking exactly as it did.
What makes this unusual is where it happens. Scanned files are often the most sensitive documents anyone owns — contracts, medical letters, identity papers — and every mainstream recognition service asks you to upload them. Here the engine runs inside your browser and the file never leaves your device.
Five languages are available: English, Hindi, Spanish, French and German. Each is a separate model of between 0.7 MB and 2.8 MB, downloaded only if you choose it, and the tool checks whether you already have one before warning you about a download. Choosing correctly matters more than it sounds — a page of Hindi read as English comes back as confident nonsense.
Accuracy depends on the scan and on the language, and both are measured rather than guessed at. On clean printed test pages the four Latin languages recover every word; Hindi recovers roughly three in five, because Devanagari is a harder script with far more ways to be slightly wrong. Recognition is limited to 50 pages at a time because it is memory-heavy, and handwriting is not read at all.
How to make a scanned PDF searchable
- 1
Open the scan
Choose the scanned PDF or drop it here. The recognition engine is prepared while you do, and the file itself never leaves your device.
- 2
Pick the language
Choose the language the document is written in — English, Hindi, Spanish, French or German — so the recogniser knows what it is reading.
- 3
Read and download
Each page is read in turn, with progress shown page by page, and the searchable PDF downloads to you.
What OCR PDF accepts and produces
- Input
- One scanned PDF, in English, Hindi, Spanish, French or German.
- Output
- The same scan with an invisible text layer over the words, so it can be searched, selected and copied while looking untouched.
- Limits
- 50 pages per run, because recognition is memory-heavy. Accuracy depends on the scan and on the language — measured recall is far lower for Hindi than for the Latin languages. Handwriting is not read.
- Offline
- Yes, once the engine and your chosen language model have been fetched. Each language is a separate download of between 0.7 MB and 2.8 MB.
Why OCR PDF here
- The scan looks exactly as it did — the recognised text is invisible
- Runs entirely on your device, with the engine served from this site
- Works offline once the engine has been stored
- Unlocks Word, Text, HTML and Highlight for documents that were only images
Where your file goes
Nowhere. This tool runs inside your browser, and the page is served with a policy that forbids it from sending your document anywhere. You can confirm that yourself: open your browser’s developer tools, watch the Network tab, and process a file.
Questions
What does making a PDF searchable actually do?
It reads the words in the scanned image and stores them in the file as an invisible text layer sitting exactly over the picture of each word. The page looks identical — your scan, unchanged, with its signatures and stamps — but a reader can now search it, select text and copy from it. Nothing visible is replaced or redrawn.
How accurate is the recognition?
It depends on the scan and on the language, and both are measured rather than guessed at. On clean printed test pages the four Latin languages recover every word we look for, at around 95% confidence. Hindi is noticeably weaker — Devanagari is a harder script, with far more ways to be slightly wrong — and recovers roughly three words in five on the same kind of page. On a faint fax, a skewed photograph or handwriting, all of them degrade sharply. The result panel reports the confidence measured on your own document, which is a better guide than any of these figures. Recognition is a reading of an image, never a transcript, so check anything that matters against the original.
Is my document uploaded to be read?
No, and this is the whole reason the tool exists here. Scanned documents are usually the most sensitive things people own — contracts, medical letters, identity papers — and every mainstream OCR service asks you to upload them. Here the recognition engine runs inside your browser and the file never leaves the device.
Why does it download a few megabytes the first time?
Because the recognition engine and your chosen language model have to be on your device to run there. Together that is roughly 4.5MB for English and rather less for the other languages; the tool tells you the size before it starts. They are fetched once from this site — never from a third-party CDN — and then stored, so every run afterwards works with no connection at all. If you have already fetched a language, the tool checks and says so rather than warning you about a download you are not about to make.
Does it work offline?
Yes, once the engine has been stored. Open the tool once while connected and the engine and model are cached; after that you can disconnect entirely and recognition still runs. The very first visit needs the network for that one download.
Which languages are supported?
English, Hindi, Spanish, French and German. Each has its own model, between 0.7MB and 2.8MB, and only the one you pick is downloaded — choosing a language never fetches the other four. The selector can only offer languages whose models are genuinely stored on this site, so it will never present one your browser cannot get. Recognising the wrong language produces confident nonsense, so choosing correctly matters more than it might seem.
Is there a limit on document size?
Fifty pages at a time. Recognition is far heavier than an ordinary PDF operation — seconds and hundreds of megabytes per page — and an unbounded job on a phone does not fail politely, it takes the tab down. Larger documents can be split first and run in parts.
Can I search a scanned PDF in other tools after this?
Yes, and that is the point. Once a document has a text layer, PDF to Word, PDF to Text, PDF to HTML and Highlight all work on it exactly as they would on a document that was born digital. Those tools can also recognise a scan directly if you ask them to.
What happens to a page with no readable text?
If nothing at all can be recognised across the pages you selected, you are told so instead of being handed a file that looks processed and is not. A page that is genuinely blank simply contributes nothing, which is correct.
Worth reading about OCR PDF
- Preparing a scanned document to sendA scan hides text you cannot see and metadata you did not add. What to check, and in what order, before a scanned document leaves your hands.
- What text recognition can and cannot readRecognition reads a crisp printed page well and a faint or skewed one badly. What affects accuracy, and how to tell whether a result can be trusted.
- What makes a PDF searchableA scan is a picture of words, not words. What a text layer is, how recognition adds one, and how to check whether a PDF you have already has one.
- PDF work for teachersAssembling worksheets, splitting a scanned pile of submissions, and keeping printing costs down. The PDF jobs that come up in a classroom, done quickly.
- PDF work for accountants and bookkeepersGetting figures out of statements, bundling invoices for filing, and meeting the size limits that tax portals impose. The PDF jobs behind a finance workflow.
- PDF work on a phone, without an appMerging, compressing and scanning on a phone usually means an app that wants a subscription. What a browser can do instead, and where the real limits are.