Formats2026-08-215 min read

What Actually Happens When You Compress a PDF

PDF123 Compress runs qpdf stream recompression and object-stream packing, not image downsampling or font subsetting. Linearization barely changes file size.

PDF123 · Updated 2026-09-19

Two PDFs can look identical on screen and differ by three orders of magnitude: 50 KB against 50 MB. The gap almost never sits in the text operators. In a twenty-page contract those operators are usually a few tens of kilobytes; everything else is image samples, a fully embedded font program, and how tightly the streams and objects are packed.

That has a consequence worth stating early: "compressing a PDF" is not one action but four. They shrink different things and charge different prices, and treating them as one thing produces contradictory experience, the file compressed but did not shrink, or it shrank but the pages went soft.

Diagram of a PDF shrunk by qpdf stream recompression and object-stream packing, with linearization as a separate non-size step

Where the size comes from

Most "why is this PDF huge?" cases start with a high-resolution scan or photo still embedded at capture size. The arithmetic is worth doing once: a letter page at 8.5 × 11 inches captured at 300 DPI is about 2550 × 3300 pixels, and at three bytes per pixel that is roughly 25 MB uncompressed. Twenty such pages often land at 60 to 80 MB before anyone touches compression.

The second source is fonts. A PDF can embed an entire font program so that any machine renders the text identically, and that cost is priced by the size of the glyph tables rather than by how much of the text is actually read. It is why font subsetting is a separate size lever.

The third is packing itself. The same content can be stored with wasteful compression settings, or its objects can sit scattered through the file, each carrying its own bookkeeping overhead. None of that changes a single pixel or glyph, and it is exactly the part that stream recompression can reach.

"Compress" is four different things

There are four levers on size, they act at different layers, and their costs barely overlap. Choosing the wrong one is the most common reason a compression run appears to do nothing.

Lever What it touches Size win Cost
Stream recompression and object streams How content streams are compressed, how objects are stored and packed Depends on how loose the original was Page appearance unchanged
Image downsampling Pixel count, that is, resolution Large Resolution permanently reduced
Lossy image re-encoding The bitmap's byte representation Large Image quality reduced
Font subsetting The embedded glyph tables Moderate Glyphs not used are no longer editable
Linearization The order of objects Near zero Buys first-page load time instead

The first four rows all get called "compression" loosely, but only the first leaves the content alone. It is also the row whose payoff is easiest to misjudge: a scan whose size is mostly raw image samples has little redundancy for it to remove, and the levers that would cut it substantially are downsampling and lossy re-encoding, at the price of losing that resolution and quality for good.

Which of the four PDF123 Compress does

Compress a PDF (POST /api/v1/misc/compress-pdf) does the first row only. It runs qpdf with compression of uncompressed streams (--compress-streams=y), a second pass over streams already compressed with Flate (--recompress-flate), and packing of objects into object streams (--object-streams=generate).

It does not downsample images, re-encode bitmaps to a lower JPEG quality, or subset fonts. Pages should look the same, and the savings come from Flate recompression and object-stream packing.

The failure condition deserves the same clarity. When a file's size is mostly image data already compressed as JPEG, Flate has nothing left to squeeze, because those image bytes sit outside what the three switches above touch. That is why a scan pack may come back only a few percent smaller while a text-heavy PDF of similar size gains noticeably more.

The inverse path is Decompress PDF, which expands streams for inspection (--qdf, object streams disabled) and likewise does not change how pages look.

When to use a different path

What you need The path to use
Identical appearance, just remove packing overhead Compress
A scan pack that is genuinely too big and quality loss is acceptable Downsampling or lossy re-encoding, not this Compress
Still editable, but the next character will not type Different export settings or source file, see font subsetting
The browser should see page one sooner Linearization solves first paint, not size
The file will not open at all Repair, a structural problem rather than a size problem

One trap sits in that last row: Linearize PDF is not a separate implementation here, it is an alias of Compress. Both point at the same operation, the canonical id is misc/compress-pdf with linearize-pdf as an alias, the portal always sends linearize=true and optimizeLevel=1 for that tool, the server reads neither field, and qpdf is never passed --linearize. The operation that does linearize a file is Protect, when it encrypts.

So if what you want is the same appearance in a smaller file, Compress is the whole answer. If you want lossy image shrink or font subsetting, this path is the wrong lever. The file comes back directly, with no account required. To see where objects, streams and the cross-reference table actually live, continue with What's inside a PDF.

Open tool
Process in the browser — no watermark, files removed after the job.
Open tool