Product2026-08-313 min read

Why Self-Host a PDF Toolkit Instead of Another Free Web Converter

Browser converters fit anonymous one-offs. Self-hosting the same catalog gives REST, MCP, a CLI, no ads in your stack, and files that never leave your network.

PDF123 · Updated 2026-09-17

Browser-only PDF converters optimize for upload, click, download. Ads often fund the page, and there is usually no stable developer surface beyond the form. That product shape is correct for a one-off human job. It is the wrong shape when you need automation, a retention policy, or files that must stay on your network.

What browser converters optimize for

A public converter sells convenience: open a tab, drop a file, get a result, leave. The UI is the product. Scripts either scrape HTML or hit undocumented endpoints that change without notice. Retention is whoever operates the temp store. For a single merge of two personal PDFs, that trade is fine.

The moment you need the same job from CI, an agent, or an air-gapped network, the missing pieces show up at once: auth, OpenAPI, idempotent retries, and a catalog that does not silently diverge from the form you clicked last week.

What you get by running this catalog yourself

The hosted site and a self-hosted instance share the same operations. Compress and Merge are the same ops whether the request came from the portal or from POST /api/v1/…. Self-hosting adds control around that shared catalog:

  • REST + OpenAPI at /v1/openapi.json: every catalog tool is a stable endpoint, not only a form.
  • MCP at /mcp: agents discover and call the same operations without inventing scrapers.
  • CLI (pdfx): local pdf-core, or --cloud against an API key on your base URL.
  • No desktop client: Docker Compose is the offline story; plugins and Windows installers are out of scope on purpose.
  • Ads stay off your stack: AdSense on the public site is not part of a self-hosted deploy. Uploads never leave the machines you run.

Compose brings up the Rust pdfx-server and the Next.js portal together so humans and scripts share one catalog. Native dev (task pdfx:dev plus the portal) uses the same ops path. Details live on Self-host.

OCR on either deployment still returns Markdown (text/markdown) from /api/v1/misc/ocr-pdf, not a re-layered searchable PDF. Relocating the server changes where bytes land; it does not invent a different op contract.

What we intentionally do not copy

Online “Edit PDF” canvases, visual compare, blank Create PDF, and consumer Sign UIs need different product work. We keep the form-based tool matrix and the API instead. Self-hosting does not unlock a second UI tree; it relocates the same tree behind your firewall.

Skipping those consumer surfaces is deliberate. A canvas editor would be a second codebase with its own feature lag relative to OpenAPI. The constraint that matters for private use is “bytes stay here,” not “there is a .exe.” For why there is no desktop fork, see Why We Skipped the Desktop App.

When hosted is enough, and when it is not

Use the public site when you want zero setup and an anonymous one-off. Catalog tools and anonymous /api/v1/ calls need no account; hosted responses advertise rate-limit headers, and caps exist so the free surface does not become someone else’s batch farm (anonymous rate limiting).

Self-host when retention rules, private networks, or private automation require it. You then own uptime, image upgrades, disk for temporary files, TLS, and key rotation (SECURITY_CUSTOMGLOBALAPIKEY for a fixed gate on your box). The operations stay the same; where the bytes land is what changes.

Try it

Start from Self-host and Developers. For the agent-facing API surface on either base URL, read Built for AI agents, not just browsers. For a longer cost/benefit split, see What self-hosting actually buys you.

Open tool
Process in the browser — no watermark, files removed after the job.
Open tool