Why is my PDF so large?
A ten-page document has no business being forty megabytes, and yet it often is. Before reaching for a compressor it is worth knowing what is actually taking up the space, because that decides how much you can get back — and whether compressing is the right move at all.
Almost always: images
In a PDF that is unexpectedly large, the bytes are nearly always pictures. A page scanned at 600 DPI in full colour is roughly 35 megapixels before compression; twenty of them make a file that no amount of clever structural work will meaningfully shrink, because the images are the file.
This is why a document that began life in a word processor and one that came off a scanner behave so differently. The first is text, lines and a handful of embedded fonts — kilobytes. The second is photographs of paper.
If your document is scanned, compression works on the images and the savings can be dramatic. If it is not, expect much less, and be suspicious of any tool that promises otherwise.
Sometimes: fonts, and the same thing stored twice
Every font used in a document is normally embedded so that it renders identically everywhere. A full CJK font can be several megabytes on its own, and a file assembled from several sources can carry the same font more than once because each source brought its own copy.
The same duplication happens with images. Merge four reports that share a letterhead logo and the logo may be stored four times. Rebuilding the file's structure removes that kind of waste without touching how anything looks — which is the part of compression that costs you nothing.
What compression will not fix
A PDF cannot be made smaller than the information it contains. If a file is large because it genuinely holds two hundred pages of photographs, the honest options are to send fewer pages or to accept the size.
Extracting the pages that matter is often the better answer, and a much better one for privacy: sending three pages of a contract rather than the whole thing means the other ninety-seven were never shared at all.
Beware of any compressor that hits an impressive number by flattening your pages into images. The file gets smaller and the text stops being text — no searching, no selecting, no copying, and no way back.
A sensible order to work in
Remove what you do not need first, then compress what is left. Compressing a hundred pages and then deleting ninety of them means most of the work was spent on pages you threw away.
If you are doing this regularly, the steps can be chained so the document is handled once instead of downloaded and re-uploaded between each one.
Working out which case you have, in one second
Open the document and try to select a sentence with the cursor. If the words highlight individually, it is a text document and its size comes from structure, fonts or a handful of oversized images. If nothing highlights, or the entire page highlights as a single block, it is a scan and the size is the images.
That one test tells you what to expect before you spend any time. Text documents usually compress modestly, because there was little waste to remove. Scans can compress dramatically, because there is a great deal.
If a document turns out to be a scan and you control the scanner, the largest saving available is not compression at all — it is scanning again at 200 or 300 DPI in greyscale rather than 600 DPI in colour. That is often a tenfold difference for a page that reads identically.