Skip to content

A black box is not redaction

Redacted court filings and contracts leak with tiresome regularity, and almost always the same way: someone drew a black rectangle over the text and saved the file. The rectangle is a shape sitting on top. The words are still underneath it, selectable, copyable, and readable by any parser.

Why the black rectangle fails

A PDF is a list of drawing instructions, not a picture. 'Place these glyphs here' and 'fill this rectangle with black' are two separate instructions, and the second does not delete the first. It is drawn afterwards, so it appears on top.

Anyone can select the text through the rectangle and paste it elsewhere. Anyone with a parser can read it without even opening the document. The information was never removed — it was hidden from one particular way of looking.

How to check a document you have been sent

Open it and try to select the text under the redaction. If your cursor highlights something, the content is still there.

A more reliable check is to extract the document's text and read what comes out. If the redacted name appears, the redaction is decorative — and if you were the sender, that document has already gone.

What real redaction has to do

The content itself has to go. In practice that means the affected pages are rebuilt so the removed material is not present in the file in any form, rather than being covered, moved or made invisible.

There is a real cost, and any tool that does not mention it is not being straight with you: a rebuilt page is an image, so its text is no longer selectable or searchable. That is the price of the content genuinely being gone. Pages you did not redact are untouched.

The part people forget: metadata

A document carries more than its pages. Author, title, the software that produced it, creation and modification times, and sometimes a block of XMP holding all of that again — none of it is visible on the page, and all of it travels with the file.

Redacting the pages of a document whose properties still name the author and the original filename is a job half done. Strip the metadata as well, and do it after the redaction rather than before.

How to check your own document

There is a test that takes ten seconds and settles it. Open the redacted PDF, select all the text on the page, copy it, and paste it into a plain text editor. If a name you redacted appears in the paste, the redaction did not work — and anybody who receives the file can run the same test.

For a scanned document there is a second copy to worry about. If the scan has been through recognition it carries an invisible text layer behind the image, so a redaction that only altered the picture leaves every word still extractable. Extract the text and check it, rather than trusting how the page looks.

Metadata is the third place content hides. A PDF can carry the author, the organisation, the software that produced it, and in a file assembled from several sources, all of that for each source. None of it appears on any page.

Where redaction leaks in practice

Almost always in one of four ways, and none of them look wrong on screen. A black rectangle drawn as an annotation sits above the text rather than replacing it. A highlight set to opaque black does the same. A page image covered with a filled shape leaves the recognised text layer beneath untouched. And a redacted page that was then merged into a bundle can reintroduce an earlier, unredacted version of itself if the wrong copy was used.

The common thread is that every one of these is a drawing operation. Redaction is a deletion operation, and the difference is invisible until somebody looks.