Skip to main content
PPDF123

Convert PDF to XML

PDF to XML imports a PDF through LibreOffice and returns an Office-flavored XML file. It is a best-effort text and layout export, not a copy of the PDF's structure or tags.

Drag and drop files here, or click to choose

Accepts PDF · up to 500.0 MB per file · up to 20 files at once

You can also paste a file with Ctrl/Cmd+V, or from the clipboard menu.

PDF to XML sends the file through LibreOffice `writer_pdf_import` with `outputFormat=xml`. The download is an Office-flavored XML document (`octet-stream`), not a reconstruction of PDF content streams or tagged-PDF structure.

This is the same LibreOffice import family as Convert to Word, aimed at XML instead of DOCX. Layout, tables, and fonts are best-effort. Scans with no text layer import as pictures or empty runs.

There are no options on this page. Password-protected files are unlocked for you on this page once you enter the password; API callers run Unlock first. OCR Markdown cannot be sent here.

If you need rows from a table, PDF to CSV is the layout heuristic. If you need words only, PDF to Text is smaller.

Features

  • LibreOffice `writer_pdf_import` to XML
  • Not a PDF-structure dump and not OCR
  • No page-range or schema options
  • Anonymous API at /api/v1/convert/pdf/xml

When to use this tool

  • Feed a born-digital PDF into a pipeline that already eats Office XML
  • Inspect how LibreOffice tokenized the document
  • Get an XML draft when DOCX is the wrong next step

How do I convert a PDF to XML?

  1. Upload a text-based PDF.
  2. Click Process. There are no format choices on this page.
  3. Download the `.xml` file.
  4. Open it in LibreOffice or an editor and treat it as a draft.

Limits and edge cases

  • Not tagged-PDF or XFDF/XFA extraction
  • Not OCR
  • Tables and pagination will not match the PDF
  • Upload limit on this website: 500 MB per file, sent in chunks above about 95 MB. A single direct API request body is capped at 100 MB.

Examples

  • A text memo often becomes a large Office XML file with paragraph runs
  • A scan-only PDF often becomes XML that mostly wraps images

Privacy for this tool

Files are uploaded over HTTPS, processed in memory or a short-lived temporary directory on our servers, and deleted when your result is ready. We do not keep copies for later browsing, training, or advertising profiles. See the Privacy Policy for retention details and AdSense cookie disclosures.

Frequently asked questions

Is this the PDF's internal XML?
No. It is LibreOffice exporting the imported document as XML, not the PDF object tree.
Why is a scan almost empty?
LibreOffice has no text to import. Use OCR to Markdown for words from pixels.
Can I choose another XML flavor?
No. The converter defaults outputFormat to xml and only allows that value on this path.
How is this different from Convert to Word?
Same importer, different output. Word gives docx/doc/odt. This path gives xml.

Last updated:

Call this from code

Every tool on this site is a plain REST endpoint - no account or API key needed for anonymous use. Built for AI agents and developers as much as for browsers.

curl -X POST "https://pdf123.xyz/api/v1/convert/pdf/xml" \
  -F "[email protected]" \
  -o output.pdf

Also available as an MCP tool for agent clients that speak Model Context Protocol (JSON-RPC 2.0 over POST /mcp). Full API reference

Use it from an AI agent

Skill

Claude Code, Codex, Cursor and other AI agents can run PDF to XML for you with this skill: /skills/pdf123-pdf-to-xml.md

All agent skills

Files are used only for this processing job and deleted automatically afterward.