Product2026-09-283 min read

Scan Pack to Text: OCR Then Compress the PDF You Still Keep

A scan pack needs two steps: OCR returns Markdown text for agents, while Compress shrinks the original PDF with stream recompression. Order and outputs differ.

PDF123 · Updated 2026-09-28

A "scan pack" is usually a stack of page images stapled into one PDF: heavy on disk, empty for search and copy. Two tools people chain together do not produce one magic searchable PDF on this site. OCR returns Markdown. Compress rewrites the PDF you still keep.

That split is the whole workflow.

What OCR actually gives you here

POST /api/v1/misc/ocr-pdf runs recognition and downloads a .md file (text/markdown): native PDF text fused with OCR from image pages. It does not write a hidden text layer back into the PDF. If you expected select-and-search inside the same file, that is a different product shape (classic OCRmyPDF-style output). This endpoint is for text an agent or pipeline can read.

Auto vs Force OCR:

  • Anything other than ocrType=force-ocr (including the portal label "Normal") maps to Auto: use existing text where present; OCR image-only pages.
  • force-ocr re-reads every page, including pages that already have a text layer.

There is no language picker on the portal form. Leftover tessdata-style fields from older paths (languages, deskew, clean flags, and similar) are ignored by the Rust ocr_pdf path. Auto vs Force, and why recognition fails on bad scans, are covered in How OCR reads a scanned PDF.

OCR does not shrink the PDF. The Markdown download is a second artifact beside the original scan pack.

Where Compress fits

Compress (POST /api/v1/misc/compress-pdf) runs qpdf stream compression: --compress-streams=y, flate recompress, and generated object streams. It does not invent a text layer, does not subset fonts, and does not promise photographic downsampling. On a scan-heavy file it may still shrink storage by packing streams harder; on an already tight file the gain can be small. Background on stream vs image levers: What actually happens when you compress a PDF.

Use Compress on the PDF you will archive or email. Use the Markdown OCR output when you need the words. Compressing first does not make OCR more accurate; OCR first does not make the PDF smaller.

A practical order

  1. If the viewer will not open the pack, Repair or Get Info first.
  2. OCR the scan pack to Markdown and spot-check a few pages of text.
  3. Separately, Compress a copy of the PDF if size is the remaining problem.
  4. Keep the untouched original when policy requires it.

Encrypting a file still needs an authorized Unlock before either step. Secrets need Sanitize / Auto Redact, not OCR. Table export that returns 204 after a scan means no text columns were found—OCR Markdown first, then decide whether a spreadsheet tool even applies (Empty table export (204)).

Outputs you should keep track of

After a typical run you may hold three artifacts: the original scan pack, a .md OCR download, and (optionally) a compressed PDF copy. Label them. Emailing only the Markdown loses the page images; emailing only the compressed PDF still has no select-and-search layer from this OCR path. Pipelines that need both words and a smaller PDF must store both outputs or re-run one side later.

Force OCR on a born-digital pack with a broken text layer can still be useful; Auto on a clean born-digital pack may skip the OCR runtime entirely. Neither mode writes text back into the PDF.

What this workflow is not

Scan packs fail when teams assume one button both extracts text and shrinks pixels the way a desktop OCRmyPDF pipeline might. Here the text leaves as Markdown; the PDF stays a PDF unless you Compress it on purpose. There is no step on this site that returns "the same PDF, now searchable and smaller" from a single call.

For HTTP automation of the same ops, see Developers.

Open tool
Process in the browser — no watermark, files removed after the job.
Open tool