Convert PDF to Text
PDF to Text extracts the text that already exists in a PDF and returns it as a plain .txt file, or as RTF. It does not run OCR, so image-only scans come out empty.
Drag and drop files here, or click to choose
Accepts PDF · up to 500.0 MB per file · up to 20 files at once
You can also paste a file with Ctrl/Cmd+V, or from the clipboard menu.
PDF to Text extracts words that already exist as text objects. The default `txt` path uses the same native extractor as other pdf-core ops (`pdfio::extract_text`) and returns `text/plain`. It does not run OCR and it does not call LibreOffice.
Choose `rtf` and the file goes through LibreOffice `writer_pdf_import` instead. That path needs `soffice` on the server and rebuilds a rough rich-text document. Expect substitution and broken tables; treat RTF as a draft.
Image-only scans yield an empty or nearly empty `.txt`. OCR on this site returns Markdown, not a searchable PDF you can send back here.
There is no page-range field. Every page is read. Password-protected files are unlocked for you on this page once you enter the password; API callers run Unlock first.
If you need headings and lists rather than a text dump, PDF to Markdown is the closer tool. If you need tables as rows, use PDF to CSV or PDF to Excel.
Features
- Default TXT via native text extraction, not pdftotext and not OCR
- Optional RTF through LibreOffice `writer_pdf_import`
- Reads every page; no page-range parameter
- Anonymous API at /api/v1/convert/pdf/text
When to use this tool
- Dump a text-layer report into grep or a script
- Get a rough RTF draft from a born-digital memo
- Check whether a PDF even has a text layer
How do I extract text from a PDF?
- Upload a PDF that already has selectable text.
- Keep TXT, or choose RTF if you need a LibreOffice rich-text draft.
- Click Process and download `.txt` or `.rtf`.
- Spot-check a few paragraphs. Scans without a text layer will look empty.
Limits and edge cases
- Not OCR; scans produce little or no text
- No page-range control
- Multi-column visual order is not reconstructed for TXT
- Upload limit on this website: 500 MB per file, sent in chunks above about 95 MB. A single direct API request body is capped at 100 MB.
Examples
- A selectable annual report usually yields a long .txt of its paragraphs
- A photographed whiteboard PDF usually yields a blank or tiny file
Privacy for this tool
Files are uploaded over HTTPS, processed in memory or a short-lived temporary directory on our servers, and deleted when your result is ready. We do not keep copies for later browsing, training, or advertising profiles. See the Privacy Policy for retention details and AdSense cookie disclosures.
Frequently asked questions
- Why is the text file empty?
- The PDF has no text objects for the extractor to read. Scans need OCR to Markdown; that Markdown cannot be sent back into this tool.
- Is TXT the same as copying from a viewer?
- Close, but not identical. Order follows how text is stored in the page content stream, which can differ from visual reading order in multi-column layouts.
- When should I pick RTF?
- Only when you want LibreOffice to rebuild a document. TXT is the direct extract and does not need soffice.
- Does this keep fonts and layout?
- TXT is plain text. RTF is a best-effort rebuild, not a pixel match.
Last updated:
Call this from code
Every tool on this site is a plain REST endpoint - no account or API key needed for anonymous use. Built for AI agents and developers as much as for browsers.
curl -X POST "https://pdf123.xyz/api/v1/convert/pdf/text" \
-F "[email protected]" \
-F "outputFormat=txt" \
-o output.pdfAlso available as an MCP tool for agent clients that speak Model Context Protocol (JSON-RPC 2.0 over POST /mcp). Full API reference
Use it from an AI agent
SkillClaude Code, Codex, Cursor and other AI agents can run PDF to RTF (Text) for you with this skill: /skills/pdf123-pdf-to-text.md
Files are used only for this processing job and deleted automatically afterward.