Skip to content

What makes a PDF searchable

Two documents can look exactly the same on screen and behave completely differently: one lets you select a sentence and search for a name, the other is a picture of a page. The difference is invisible until you need it, and it is decided by how the file was made rather than by how it looks.

A PDF page is instructions, not a picture — usually

A PDF page is normally a list of drawing instructions: put these characters, in this font, at these coordinates. Because the characters are really there, software can select them, search them, copy them and read them aloud.

A page can equally well contain a single instruction that says 'draw this photograph here'. It looks like a document, prints like a document, and contains no text at all — the words are shapes in an image, exactly as they are in a photograph of a road sign.

The quickest way to tell which you have is to try to select a line. If the cursor sweeps across the words without highlighting anything, or highlights a rectangle covering the whole page, there is no text layer.

How documents end up as pictures

Scanning is the obvious way. A scanner produces an image, and unless something adds a text layer afterwards, an image is what the PDF holds. The same applies to a photograph of a page taken on a phone.

Less obviously, several common operations deliberately convert pages to images. Genuine redaction does it, because that is the only way to be certain removed words are gone rather than merely covered. Inverting a document's colours does it, because the colours in a PDF come from too many places to rewrite reliably. In both cases the loss of the text layer is the price of the operation working at all, and a tool that does either without telling you is worth distrusting.

There is also a category of converter that produces image-only PDFs by accident: anything that renders a page and screenshots it. If a tool advertises perfect visual fidelity and produces a file whose text cannot be selected, that is what it did.

Adding text back to a picture

Recognition reads the shapes in an image and works out what letters they are, then writes those letters into the page invisibly, positioned behind the picture. The page looks unchanged and becomes searchable.

It is a reading rather than a transcript, and the accuracy depends on the scan: resolution, how straight and sharp the page is, the typeface, and the language. A crisp scan of printed text reads very well; a faint fax or a photograph taken at an angle reads badly, and handwriting is not read at all.

That matters for what you do next. Anything that works on the text of a document — searching, extraction, summarising, scanning for personal data — is working on the recognition's output, not on the document. On a poor scan it will miss things, and it will miss them silently.

Starting from text instead

If you are producing the document rather than receiving it, the whole problem is avoidable: a PDF written from text is searchable from the start, with no recognition step and no accuracy to worry about.

That holds whether the source is a Word document, a Markdown file or text typed straight into a box. What matters is that the words are placed as characters rather than drawn as an image — which is also why the resulting file is a fraction of the size of a scan, and why it stays sharp at any zoom.

It is worth checking the output of any converter once, in the way described above. A tool that produces a picture of your text has given you something that looks right and cannot be used.