Skip to content
Runs on your deviceWorks offline

Find and remove personal data from a PDF

Find personal data in a PDF automatically, then remove what you choose.

Your device

Or drop one anywhere on this page

Local processingNo server uploadYour files are processed in your browser and are not uploaded to our servers.
Local processingNo server uploadYour files are processed in your browser and are not uploaded to our servers.

What to know about Auto-Redact PII

Reading a long document line by line looking for an account number is how personal data gets missed. This scans every page for the patterns that identify people — Aadhaar and PAN numbers, payment cards, bank and IFSC details, emails, phone numbers — shows you what it found and where, and removes only what you tick. The scan and the removal both happen on your own device.

Most personal data in documents is not leaked by carelessness with the obvious things. It is leaked by length. A forty-page agreement with an annexure of bank details, an HR file with a scanned identity page stapled to the back, an invoice batch exported in one PDF — the sensitive part is real, it is somewhere, and nobody is going to read every line before sending it. Automated scanning is worth doing because attention does not scale and a text search does.

What the scan actually does is match formats. Aadhaar numbers, PAN, GSTIN, payment cards, bank accounts, IFSC codes, passport numbers, emails, phone numbers and IP addresses all have shapes that can be recognised, and several of them carry a check digit that can be verified — the Verhoeff digit on an Aadhaar number, the Luhn digit on a payment card, the mod-36 character that closes a GSTIN. A value whose checksum verifies is almost certainly the real thing and is graded critical. A value that has the right shape but fails the check is still shown, one grade lower, because a mistyped or badly-scanned Aadhaar number is exactly as personal as a correct one.

That grading reflects a deliberate bias. Detection here errs towards showing you too much rather than too little. Bank account numbers, in particular, have no standard format anywhere in the world, so any long run of digits might be one — and flagging a purchase order number by mistake costs you a second's glance, while missing an account number defeats the point of running the scan. Findings are sorted so the ones that matter most are impossible to miss, and you can hide the lower grades once you have dealt with them.

Two limits are worth being blunt about. The first is that this cannot see a picture. Pattern matching needs text, and a scanned page is an image; a photographed Aadhaar card in a PDF is invisible to the scan, which will happily report that it found nothing. Pages with no text layer are counted and named for exactly this reason, and the fix is to run OCR first and scan the searchable copy. The second is that formats are not meaning. A person's name in a sentence, a home address, a date of birth in prose, a medical detail — all personal, none of them detectable by pattern. This is a very good first pass over a document nobody has time to read closely. It is not a guarantee that a document is clean.

Nothing is removed unless you choose it. Every finding starts unticked, and the removal only runs when you confirm it — the same rebuild-from-pixels removal as manual redaction, so the selected content is genuinely gone from the file rather than hidden under a rectangle that any parser can see straight through. Pages you did not select keep their text, links and structure exactly as they were. And because a document's properties survive anything you do to its pages, the scan also lists the author, title and creator fields that are set, so a file whose visible details you have just cleaned does not go out still carrying a name in its metadata.

How to find and remove personal data from a PDF

  1. 1

    Open the document

    Pick the PDF or drop it here. It is read on your device, and no copy of it is sent anywhere.

  2. 2

    Let it scan every page

    Each page's text is checked against the patterns that identify people. You get a list grouped by how serious each finding is, with the page it sits on.

  3. 3

    Choose what should go

    Every finding is shown masked, never in full. Tick the ones that are genuinely personal and leave anything the scan misread.

  4. 4

    Remove and download

    The pages carrying your selections are rebuilt without them, and the cleaned file downloads straight to you.

What Auto-Redact PII accepts and produces

Input
One PDF that contains real text. Scans have no text layer until OCR has been run on them.
Output
A list of what was found, with each item's page and position — then a PDF with the items you selected genuinely removed.
Limits
Detection reads a text layer, so it cannot see personal data inside a photograph or an un-OCRed scan. It recognises formats, not meaning: a name in prose is not detected, and something shaped like an account number may be flagged when it is a reference. Review the findings; they are a starting point, not a verdict.
Offline
Yes, once the rendering engine has been fetched on your first visit.

Why Auto-Redact PII here

  • Every page is checked, including the appendices nobody re-reads
  • Findings are graded, so an Aadhaar number is not buried among phone numbers
  • Values are shown masked, so the list itself cannot leak the document
  • Checksums confirm Aadhaar, PAN, GSTIN and card numbers rather than guessing by shape
  • Nothing is removed until you select it — the tool never edits on its own judgement
  • Selected content is genuinely deleted from the file, not covered with a black box

Where your file goes

Nowhere. This tool runs inside your browser, and the page is served with a policy that forbids it from sending your document anywhere. You can confirm that yourself: open your browser’s developer tools, watch the Network tab, and process a file.

Questions

What kinds of personal data does the scan look for?

Aadhaar numbers, PAN, GSTIN, payment card numbers, bank account numbers and IFSC codes, passport numbers, email addresses, phone numbers and IP addresses. Where a format carries a check digit — Aadhaar, cards, GSTIN — that digit is verified, so a confirmed match is graded higher than something that merely has the right shape.

Will it find personal data in a scanned document?

No, and this is the most important limitation to understand. Detection reads the text layer of a PDF, and a scan is a picture of a page with no text layer at all. A photographed identity card is completely invisible to the scan, which will report finding nothing. Run the document through OCR first and scan the searchable version — the tool counts pages with no text and tells you when this applies.

Does it remove anything automatically?

Never. The scan only reads; every finding starts unselected and nothing is changed until you choose items and confirm. That is deliberate — pattern matching has false positives, and a tool that silently deleted an invoice number it mistook for an account number would be corrupting documents rather than protecting them.

Why does the list show numbers with dots through them?

Because a list of every Aadhaar number in your document is itself a disclosure. It appears on a shared screen, in a screenshot, in a support ticket. Enough of each value is shown for you to recognise which one it is, and the rest is masked. Nothing in the interface ever holds the full values.

How is this different from Redact PDF?

Redact PDF is for content you have already spotted: you drag a box over it. This is for content you have not spotted — it reads the whole document and tells you where to look. Both use the same removal engine, so once you have chosen, what happens to the file is identical.

Is the found text really deleted from the file?

Yes. Pages carrying a selected finding are rendered to pixels with the selections covered, and that image replaces the page's contents. There is no text object left behind for a parser to read back, which is the flaw in drawing a black rectangle over something and saving it.

Can it miss something?

Yes, and you should assume it can. It matches formats, so it does not recognise a person's name in a sentence, a handwritten note, an address, or an identifier in a layout it has no pattern for. Treat the findings as a thorough first pass over a document too long to read closely, then check anything that matters yourself.

Why is something flagged that is not personal data?

Because bank account numbers have no standard format, so any run of nine to eighteen digits could be one. The tool deliberately errs towards showing you too much: a false alarm costs you a moment, while a missed identity number is the failure the tool exists to prevent. Findings whose check digit does not verify are graded lower so they are easy to skip.

Does it check the document properties as well?

It reports them. An author name in the document properties survives every redaction you make on the page, so the scan lists the properties that are set and sends you to Remove PDF Metadata to strip them. That is a separate operation because it rewrites the file in a different way.

Is my document uploaded to be scanned?

It is not. Reading the text, matching the patterns and rebuilding the pages all happen in a worker inside this tab. Open your browser's network panel and scan a file — nothing carrying the document leaves. Turn the network off completely and it still works.

Worth reading about Auto-Redact PII