What Actually Happens When You Compress a PDF
PDF123 Compress runs qpdf stream recompression and object-stream packing, not image downsampling or font subsetting. Linearization barely changes file size.

Two PDFs can look identical on screen and differ by three orders of magnitude: 50 KB against 50 MB. The gap almost never sits in the text operators. In a twenty-page contract those operators are usually a few tens of kilobytes; everything else is image samples, a fully embedded font program, and how tightly the streams and objects are packed.
That has a consequence worth stating early: "compressing a PDF" is not one action but four. They shrink different things and charge different prices, and treating them as one thing produces contradictory experience, the file compressed but did not shrink, or it shrank but the pages went soft.
Where the size comes from
Most "why is this PDF huge?" cases start with a high-resolution scan or photo still embedded at capture size. The arithmetic is worth doing once: a letter page at 8.5 × 11 inches captured at 300 DPI is about 2550 × 3300 pixels, and at three bytes per pixel that is roughly 25 MB uncompressed. Twenty such pages often land at 60 to 80 MB before anyone touches compression.
The second source is fonts. A PDF can embed an entire font program so that any machine renders the text identically, and that cost is priced by the size of the glyph tables rather than by how much of the text is actually read. It is why font subsetting is a separate size lever.
The third is packing itself. The same content can be stored with wasteful compression settings, or its objects can sit scattered through the file, each carrying its own bookkeeping overhead. None of that changes a single pixel or glyph, and it is exactly the part that stream recompression can reach.
"Compress" is four different things
There are four levers on size, they act at different layers, and their costs barely overlap. Choosing the wrong one is the most common reason a compression run appears to do nothing.
| Lever | What it touches | Size win | Cost |
|---|---|---|---|
| Stream recompression and object streams | How content streams are compressed, how objects are stored and packed | Depends on how loose the original was | Page appearance unchanged |
| Image downsampling | Pixel count, that is, resolution | Large | Resolution permanently reduced |
| Lossy image re-encoding | The bitmap's byte representation | Large | Image quality reduced |
| Font subsetting | The embedded glyph tables | Moderate | Glyphs not used are no longer editable |
| Linearization | The order of objects | Near zero | Buys first-page load time instead |
The first four rows all get called "compression" loosely, but only the first leaves the content alone. It is also the row whose payoff is easiest to misjudge: a scan whose size is mostly raw image samples has little redundancy for it to remove, and the levers that would cut it substantially are downsampling and lossy re-encoding, at the price of losing that resolution and quality for good.
Which of the four PDF123 Compress does
Compress a PDF (POST /api/v1/misc/compress-pdf) does the first row only. It runs qpdf with compression of uncompressed streams (--compress-streams=y), a second pass over streams already compressed with Flate (--recompress-flate), and packing of objects into object streams (--object-streams=generate).
It does not downsample images, re-encode bitmaps to a lower JPEG quality, or subset fonts. Pages should look the same, and the savings come from Flate recompression and object-stream packing.
The failure condition deserves the same clarity. When a file's size is mostly image data already compressed as JPEG, Flate has nothing left to squeeze, because those image bytes sit outside what the three switches above touch. That is why a scan pack may come back only a few percent smaller while a text-heavy PDF of similar size gains noticeably more.
The inverse path is Decompress PDF, which expands streams for inspection (--qdf, object streams disabled) and likewise does not change how pages look.
When to use a different path
| What you need | The path to use |
|---|---|
| Identical appearance, just remove packing overhead | Compress |
| A scan pack that is genuinely too big and quality loss is acceptable | Downsampling or lossy re-encoding, not this Compress |
| Still editable, but the next character will not type | Different export settings or source file, see font subsetting |
| The browser should see page one sooner | Linearization solves first paint, not size |
| The file will not open at all | Repair, a structural problem rather than a size problem |
One trap sits in that last row: Linearize PDF is not a separate implementation here, it is an alias of Compress. Both point at the same operation, the canonical id is misc/compress-pdf with linearize-pdf as an alias, the portal always sends linearize=true and optimizeLevel=1 for that tool, the server reads neither field, and qpdf is never passed --linearize. The operation that does linearize a file is Protect, when it encrypts.
So if what you want is the same appearance in a smaller file, Compress is the whole answer. If you want lossy image shrink or font subsetting, this path is the wrong lever. The file comes back directly, with no account required. To see where objects, streams and the cross-reference table actually live, continue with What's inside a PDF.