Skip to content

What text recognition can and cannot read

Recognition is a reading of a photograph, not a transcript. The same tool that recovers a crisp printed page almost perfectly will produce confident nonsense from a faint fax, and knowing which of those you have saves a great deal of time.

Resolution matters, up to a point

Below about 200 DPI, letters stop having enough pixels to be told apart, and the characters that suffer first are the ones that already look alike: 8 and B, 1 and l, 0 and O, rn and m.

Above roughly 300 DPI there is little further gain and a great deal more memory used. If you control the scanner, 300 DPI in greyscale is the sweet spot for text — colour rarely helps recognition and multiplies the file size.

Skew, contrast and the state of the paper

A page scanned at a slight angle is much harder to read than a straight one, because lines of text no longer sit along rows of pixels. A photograph of a page taken by hand is usually both skewed and unevenly lit.

Faint photocopies, coffee marks, staple shadows and show-through from the reverse side all add marks that look like punctuation. Recognition drops the least confident of these, which is why a poor scan often comes back with words simply missing rather than obviously wrong.

Language, and why choosing correctly matters more than it seems

A recogniser reads with a model of one language. Point an English model at a page of Hindi and it will not fail — it will return confident nonsense, because it is doing exactly what it was asked to do with a page it cannot interpret.

Accuracy also varies genuinely between languages. On the same clean test pages, the Latin-script languages recover essentially every word, while Devanagari recovers roughly three in five: conjuncts and matras give far more ways to be slightly wrong than a Latin alphabet does. That is a property of the script, not a bug, and it is worth knowing before relying on the result.

What it will not do

Handwriting is not read. The models are trained on printed text, and cursive in particular is a different problem entirely.

Recognition also does not repair a document. It adds an invisible layer of text over the picture of each word, so the scan still looks exactly as it did — which is the point, because that scan may carry a signature or a stamp worth keeping.

Once a document has been recognised it becomes useful to everything that needs real text: converting it to Word, pulling the text out, highlighting a passage, or comparing two versions of it.

What actually improves accuracy

Resolution first. Recognition works from the shape of each character, and below roughly 200 DPI the shapes stop being distinct — an 'rn' becomes an 'm', an 'l' becomes a '1'. Three hundred DPI is the number scanners settled on for text and it is the right target.

Then contrast and evenness. A page photographed under a desk lamp has a bright side and a dark side, and recognition reads the dark side badly. Getting even light on the page, or capturing it with a mode that brings the paper back to white, does more for accuracy than any setting.

Then squareness. Text that runs at an angle across the image is read line by line as if it were straight, and a few degrees of rotation is enough to merge adjacent lines. Straightening the page before recognition is worth more than re-running it afterwards.

And the language. Choosing the wrong one does not degrade the result gracefully — it returns confident nonsense, because the recogniser is matching against the wrong alphabet and the wrong dictionary.