From Markdown to PDF and Back: What Format Conversion Preserves
Markdown↔PDF keeps headings, lists, and simple tables better than pixels. Exact fonts, pagination, and layout do not survive as a lossless round trip.

Markdown and PDF optimize for different contracts. Markdown is version-control-friendly structure. PDF is a fixed page description. Converting one way and back is useful; it is not a lossless archive cycle.
Markdown to PDF: structure becomes pages
Markdown to PDF (POST /api/v1/convert/markdown/pdf) accepts a .md file or a ZIP that contains Markdown. Non-Markdown, non-ZIP uploads are rejected. UTF-8 is required for plain .md input. The pipeline renders Markdown to HTML with tables, strikethrough, and task lists enabled, wraps that HTML in print-oriented CSS (including multi-script font stacks and RTL when Arabic is detected), then prints to PDF with WeasyPrint.
What tends to survive the forward pass:
- Headings, paragraphs, lists, and simple Markdown tables
- Multi-script text the HTML path can lay out (CJK, Latin, and other scripts covered by the print CSS)
- A printable, shareable PDF for docs-as-code workflows
What does not:
- Your editor's theme fonts as a guaranteed match in every viewer
- Exact screen line breaks from the
.mdfile - Interactive Markdown features that have no PDF equivalent (live checkboxes, collapsible sections, wiki links)
A ZIP input is for bundling Markdown with assets the HTML path can resolve beside the document. Treat it as a packaging convenience, not as a guarantee that every relative image or CSS reference will look the same in every viewer.
PDF to Markdown: text layer in, Markdown out
PDF to Markdown (POST /api/v1/convert/pdf/markdown) extracts textual content into Markdown via pdf-inspector. The response is text/markdown downloaded as a .md file. Born-digital PDFs with a clean text layer convert best. Scanned pages with no text layer yield empty or useless Markdown until you OCR first—and OCR on this site also returns Markdown, not a searchable PDF layer.
Expect:
- Headings and paragraphs when font-size heuristics fire correctly (you may still renumber
#levels by hand) - Tables as Markdown tables when recognition works; multi-column layouts may linearize into a different reading order
- Running headers and footers repeated per page (trim in post)
- Images and equations-as-pictures not automatically becoming local assets or LaTeX
If you need plain prose without heading heuristics, PDF to Text is the flatter extract. Table-shaped exports that return HTTP 204 mean no columnar blocks were detected; see Empty table export (204).
What a round trip actually proves
| Survives reasonably | Usually does not |
|---|---|
| Words, heading hierarchy, list structure | Pixel-perfect page geometry |
| Simple tables as text | Exact fonts and kerning |
| A working draft for Git or agents | Print-identical pagination |
Round-tripping twice will drift. Markdown→PDF reflows through WeasyPrint; PDF→Markdown rebuilds structure from extraction heuristics. Neither step stores a lossless intermediate of the other format's layout model.
Pick a canonical direction
If you need the print layout as the system of record, keep the PDF and treat Markdown as an export or agent feed. If you need diffs, code review, and agent ingestion, prefer Markdown and treat PDF as the publish step. Do not store both as equal "sources of truth" and expect them to stay identical after edits.
A practical docs-as-code loop: edit Markdown in Git → Markdown to PDF for the shareable artifact → avoid PDF→Markdown→PDF as a daily habit. Use PDF→Markdown when you inherited a born-digital PDF and need text for an agent, then treat that Markdown as a new draft, not as a guarantee of the original pagination.
Encrypted files need an authorized Unlock before either conversion. Damaged files that will not parse should go through Get Info / Repair first. For the same endpoints over HTTP, see Developers.