OCR a Scanned PDF to Markdown
OCR to Markdown reads a scanned or image-based PDF and returns its text as a Markdown file. It does not add a text layer to the PDF, and the original file is unchanged.
Drag and drop files here, or click to choose
Accepts PDF · up to 500.0 MB per file · up to 20 files at once
You can also paste a file with Ctrl/Cmd+V, or from the clipboard menu.
OCR reads a scanned or image-based PDF and returns Markdown (`text/markdown`, filename `.md`). It does not write a hidden text layer back into the PDF, and it does not clean, deskew, or rewrite the scan.
The server fuses native PDF text with optical recognition on image pages. Force OCR (`ocrType=force-ocr`, the portal default) re-recognizes every page. Normal (anything other than `force-ocr`) uses existing text where the file already has it and OCRs image-only pages. On a clean born-digital PDF, Auto does not load the OCR runtime.
Accuracy follows scan quality, contrast, and skew. The engine has no language parameter. Leftover Java OCRmyPDF fields (`languages`, `deskew`, `clean`, `sidecar`) are ignored if a client still sends them.
This is not a preprocessor for Convert to Word or PDF to Text. Those tools still need text inside a PDF. OCR here is a parallel extract: you get Markdown, the original PDF is unchanged.
Do not OCR documents you are not allowed to digitize or redistribute. Copyright and workplace rules apply to the extracted text.
Libraries that want a searchable PDF will not get one here. The searchable artifact is the Markdown file. Open it in an editor and spot-check names and numbers that must be exact.
Crop and straighten phone captures before upload when you can. A trapezoid photo of a page wastes recognition on background. OCR is not translation; it recognizes characters, it does not rewrite the document into another language.
Jobs run server-side. Download promptly. We do not keep an OCR library of your files. Password-protected files are unlocked for you on this page once you enter the password; API callers run Unlock first.
Features
- Returns Markdown, not an OCR'd PDF
- Force OCR re-reads every page; Normal/Auto uses native text when present
- Original PDF is unchanged
- Anonymous API at /api/v1/misc/ocr-pdf
When to use this tool
- Extract clause text from archive scans
- Feed scans to agents or pipelines that want Markdown
- Recover text when a PDF is image-only
How do I OCR a scanned PDF?
- Upload a scanned or camera-captured PDF.
- Choose Force OCR (default) or Normal/Auto.
- Click Process and download the `.md` file.
- Spot-check names and numbers in the Markdown before you rely on it.
Limits and edge cases
- Accuracy varies with image quality
- Very large page counts take longer
- No language parameter on the current engine
- Not a local OCR library or offline SDK
- Upload limit on this website: 500 MB per file, sent in chunks above about 95 MB. A single direct API request body is capped at 100 MB.
Examples
- A 20-page grayscale contract scan becomes a Markdown file of clause text
- A skewed phone photo may need re-capture before OCR is useful
Privacy for this tool
Files are uploaded over HTTPS, processed in memory or a short-lived temporary directory on our servers, and deleted when your result is ready. We do not keep copies for later browsing, training, or advertising profiles. See the Privacy Policy for retention details and AdSense cookie disclosures.
Frequently asked questions
- Will OCR change how pages look?
- It does not return a PDF. The original file is unchanged; you get a separate Markdown download.
- Is this an OCR library I can embed?
- No. It is a hosted HTTP job at /api/v1/misc/ocr-pdf. See /developers. It is not a linkable SDK.
- Can I Convert to Word after OCR?
- Not as a PDF pipeline. OCR output is Markdown. Convert to Word still needs a text-layer PDF.
- Do I need Force OCR on a born-digital PDF?
- Usually no. Normal/Auto uses the existing text and skips the OCR runtime. Use Force if that layer is broken or you do not trust it.
Last updated:
Call this from code
Every tool on this site is a plain REST endpoint - no account or API key needed for anonymous use. Built for AI agents and developers as much as for browsers.
curl -X POST "https://pdf123.xyz/api/v1/misc/ocr-pdf" \
-F "[email protected]" \
-F "ocrType=force-ocr" \
-o output.pdfAlso available as an MCP tool for agent clients that speak Model Context Protocol (JSON-RPC 2.0 over POST /mcp). Full API reference
Use it from an AI agent
SkillClaude Code, Codex, Cursor and other AI agents can run OCR to Markdown for you with this skill: /skills/pdf123-ocr.md
Files are used only for this processing job and deleted automatically afterward.