How OCR Reads a Scanned PDF — and Why This Tool Returns Markdown
A scanned PDF is a picture of text. This OCR endpoint returns Markdown, not a PDF with a hidden layer. Auto uses existing text; Force OCR re-reads every page.

A scanned PDF stores pages as raster images—photographs of text, not selectable characters. OCR (optical character recognition) looks at those pixels, guesses which shapes are characters, and emits real text. On PDF123 that runs as a free browser tool and as POST /api/v1/misc/ocr-pdf; the response is Markdown, not a rewritten PDF.
The scan is still a picture
OCR does not clean, sharpen, or replace the bitmap. A crisp scan that OCR misreads still looks crisp. The image was never the quality bottleneck—the recognition was: swap the recognition engine on the same batch of scans and accuracy can shift by a wide margin, with the pixels themselves untouched.
Classic PDF OCR writes a second, invisible text layer aligned under the image so select and search land on the right words. That is what the diagram shows, and plenty of desktop tools still work that way.
What this endpoint actually returns
The OCR route calls /api/v1/misc/ocr-pdf, which runs on a standalone Rust crate, pdf-inspector (built with its ocr feature). It does not hand the whole PDF to a recognition model. pdf-inspector first classifies each page as either already having extractable native text or being image-only. Pages with native text keep that text as-is and never touch the recognition model; image-only pages go through the recognition pipeline instead—PDFium (the same open-source renderer Chrome uses internally) rasterizes the page, and ONNX Runtime runs the PP-OCRv6 Small model to read it. Both kinds of page get stitched together in order into one Markdown document, which is why the response is text/markdown and the download ends in .md.
That contract is new to this rewrite. The Java/Spring Boot backend this project used to run called OCRmyPDF, the classic "image plus hidden text layer" approach, and returned a PDF. The Rust rewrite replaced that path wholesale with pdf-inspector's selective OCR, and the output moved from PDF to Markdown along with it. The old Java backend has since been removed from the repository; it is not a toggle you can flip back to.
If you need a searchable PDF (image plus text layer), this endpoint cannot give you one. It gives you text an agent or a data pipeline can consume directly.
Auto vs Force OCR
The form has an ocrType field:
- Normal (anything other than
force-ocr) maps to Auto: use existing text where the file already has it; run OCR on image-only pages. On a clean born-digital PDF, Auto finds every page already has native text, so the request never touches PDFium or ONNX Runtime at all. - Force OCR (
ocrType=force-ocr) re-runs recognition on every page, including pages that already have a text layer. You pay time, not extra accuracy, on genuine image-only pages; you get a pass that ignores a broken or partial existing layer.
The portal default is Force OCR. There is no language control: leftover language parameters from the old Java path are ignored by the server.
The Auto/Force split happens inside the request, and neither the portal nor the API's capability probe checks it for you ahead of time. GET /api/v1/settings/get-endpoints-status always reports ocr-pdf as enabled, without actually verifying that PDFium, ONNX Runtime, or the PP-OCRv6 model are installed. The real check happens the moment a request triggers recognition: if any of them is missing, the endpoint returns 503, not a degraded Markdown that quietly falls back to native text only. The PP-OCRv6 Small model itself is about 31 MB; unless it was pre-seeded into the local cache at deploy time, it downloads over the network the first time a page genuinely needs recognition. A machine that is offline with no pre-seeded model will most likely fail right there on its first Force request rather than just run slowly.
Recognition is probabilistic
Engines match pixel patterns to statistical models of letters and word shapes. They do well on clean, high-contrast, standard-font print, and noticeably worse on:
- Low-resolution or low-contrast scans (fax-quality vs a 300 DPI original)
- Handwriting, decorative fonts, odd layouts (multi-column tables, rotated text, dense forms)
- Skewed pages that warp glyph shapes just enough to confuse matching
A blurry fax will not OCR like a clean scan, regardless of Auto vs Force. That is a limit of the model itself, not something a parameter can fix.
What OCR does not do
OCR gives you text. It does not rebuild paragraphs, headings, tables, or a reflowable Word document. Turning recognition output into editable layout is a conversion problem layered on top of recognition error, not the same job as reading a scan.
OCR a PDF runs server-side with no account. Download is Markdown. If an existing text layer looks wrong, use Force OCR so Auto does not trust that layer; to check whether a given machine actually has PDFium and the model installed, running a real request beats reading the capability probe, since the probe only confirms the endpoint exists, not that its runtime is ready. For automation keys and OpenAPI, see Developers.