Extract text from a PDF
Pull the plain text out of a PDF for editing or searching.
Or drop one anywhere on this page
What to know about PDF to Text
Get the text content out of a PDF so you can search, quote or reuse it. The text layer is read directly on your device and returned as plain text, page by page.
One thing is worth knowing before you start: this reads text that is already in the file. A scanned document has none — it is a photograph of words, and it will come back empty. That is what OCR PDF is for, and running a scan through it first gives this tool something to read.
Because a PDF records placed glyphs rather than paragraphs, reading order has to be reconstructed from where the characters sit on the page. Ordinary documents come out in the order you would read them. A heavily designed page — multiple columns, sidebars, text wrapped around figures — can come out in an order you did not expect, and no extractor can reliably do better, because the information was never stored.
The text is returned page by page, which makes it straightforward to quote a passage and still know where it came from.
How to extract text from a PDF
- 1
Choose a PDF
Drop the file onto the page.
- 2
Extract the text
The text layer is read locally and shown page by page so you can check it.
- 3
Copy or download
Copy any section to your clipboard, or download everything as a .txt file.
What PDF to Text accepts and produces
- Input
- One PDF that already carries a text layer of its own.
- Output
- Plain text, page by page, exactly as the document stores it.
- Limits
- A scanned PDF has no text to extract — it is a picture of words. Run OCR PDF over it first. Reading order is reconstructed from where the glyphs sit, so complex multi-column layouts can come out in an unexpected order.
- Offline
- Yes, once the rendering engine has been fetched on your first visit.
Why PDF to Text here
- Text is read on your device, never uploaded
- Output is organised page by page
- Copy a section or download the whole document
- Useful for searching long reports and contracts
Where your file goes
Nowhere. This tool runs inside your browser, and the page is served with a policy that forbids it from sending your document anywhere. You can confirm that yourself: open your browser’s developer tools, watch the Network tab, and process a file.
Questions
Why does my scanned PDF return no text?
A scan is a picture of a page — there is no text layer to read. Turning that image into text requires OCR, which recognises characters visually. OCR is planned but is not part of this release, so this tool returns nothing for pure scans.
Will the layout be preserved?
Partly. Reading order and line breaks are preserved as faithfully as the document allows, but multi-column layouts and tables can come out in an unexpected order, because a PDF stores positioned text rather than document structure.
Can I extract text from a password-protected PDF?
You will need to unlock it first with the password. Use the Unlock PDF tool, then extract text from the unlocked copy.
Can I copy the text from just one page?
Yes. Each page is shown separately with its own copy button, so you can take one section without the rest.
Are tables preserved?
Not as tables. A PDF stores positioned characters rather than rows and columns, so a table comes out as text in reading order. Simple tables often survive readably; complex ones do not.
What format is the downloaded file?
A plain .txt file encoded as UTF-8, with each page preceded by a marker showing its number. It opens in any text editor.
Will it read text inside images?
No. Text that is part of an image is not text as far as the file is concerned — extracting it needs OCR, which recognises characters visually. That is planned but not part of this release.
Is there a page limit?
No. Long documents take longer to read, and everything runs on your device, so a several-hundred-page report is fine on a laptop and slower on an older phone.
Worth reading about PDF to Text
- A black box is not redactionA black rectangle drawn over a name leaves the name in the file. Learn how PDF redaction really works, how leaks happen, and how to check your own document.
- What text recognition can and cannot readRecognition reads a crisp printed page well and a faint or skewed one badly. What affects accuracy, and how to tell whether a result can be trusted.
- What makes a PDF searchableA scan is a picture of words, not words. What a text layer is, how recognition adds one, and how to check whether a PDF you have already has one.
- When a PDF will not openA PDF that will not open is usually truncated, password-protected or genuinely damaged. How to tell which of the three you have, and what fixes each.