Product2026-09-083 min read

Merging PDFs Without Flattening Everything Into Images

A proper PDF merge concatenates page streams: fonts, vectors, and bookmarks stay intact. Rasterizing every page into images is a different, lossier job.

PDF123 · Updated 2026-09-17

“Merge these five PDFs” can mean two pipelines. One concatenates existing page objects into a single file. The other prints every page to a bitmap and packs the images into a new PDF. Both produce one attachment. Only the first keeps selectable text and sharp vectors.

What a structural merge does

On PDF123, Merge combines two or more PDFs by joining each file’s page objects. The server path runs qpdf with an empty base and --pages over the inputs: page streams move as pages, not as screenshots. Embedded fonts, vectors, and existing bookmarks travel with those pages. Nothing in that path rasterizes a page into an image, so a born-digital contract can still search and copy after the join.

Upload order is output order by default. The HTTP form accepts repeated fileInput parts in one request; batch size is bounded by upload limits, not by a hard two-file cap. Optional sortType=byFileName reorders inputs by filename before the join. Bookmarks from each source can survive; a merge does not invent a perfect outline tree if the inputs never had one.

A single-file upload is a no-op pass-through of that PDF. Two or more files become one merged.pdf whose pages are the concatenation of the inputs in the chosen order. The same op id is what OpenAPI, MCP, and pdfx merge expose, so a script and a browser click do not secretly disagree about “merge.”

When people accidentally rasterize

Some “combine” workflows export to images first (fax gateways, scan apps, print-to-PDF at screen resolution) or flatten annotations into bitmaps before joining. File size jumps, OCR may be required again, and zoom shows pixel edges.

Prefer a stream merge when sources are already digital PDFs. Prefer OCR only when the inputs are scans without usable text, and remember OCR here returns Markdown (text/markdown), not a re-layered PDF with a hidden text layer. Rasterizing then OCR-ing is a different product path from merging born-digital files.

If you need images of pages, that is a convert job (PDF to images), not a merge. Mixing those intents is how “merged” files become multi-megabyte photo albums of documents. Compress after a structural merge can still shrink streams (Compress); it does not undo a raster flatten you already baked in.

What merge does not promise

Merge does not normalize page sizes, unify fonts across files, or rebuild a table of contents from scratch. It does not decrypt encrypted inputs for you; unlock first if the source is password-protected. It does not replace a portfolio/package container format that some court portals expect; it produces one ordinary multi-page PDF.

Those limits are the same whether you use the browser, REST, MCP, or pdfx. The op is intentionally narrow so the output stays predictable for automation. If you need ordered multi-step work (merge, then watermark, then compress), use POST /api/v1/pipeline with an ordered steps list instead of inventing a second “smart merge” flag.

How to run it

Merge PDFs in the browser, or POST /api/v1/general/merge-pdfs with repeated fileInput parts (Developers). From a shell:

pdfx merge a.pdf b.pdf -o merged.pdf

Or against a hosted or self-hosted base:

pdfx --cloud --api-base "$API_BASE" --api-key "$KEY" merge a.pdf b.pdf -o merged.pdf

Anonymous catalog calls need no key on the public site; keys matter for stable automation identity and for gated self-hosted servers. For the same op through curl and MCP as well, see Same operation, four clients.

Open tool
Process in the browser — no watermark, files removed after the job.
Open tool