Skip to main content
PPDF123

Convert PDF to HTML

PDF to HTML converts a PDF with Poppler's pdftohtml and delivers a ZIP holding the HTML pages and their images. It is a layout dump rather than clean web markup, and it does not OCR.

Drag and drop files here, or click to choose

Accepts PDF · up to 500.0 MB per file · up to 20 files at once

You can also paste a file with Ctrl/Cmd+V, or from the clipboard menu.

PDF to HTML runs Poppler `pdftohtml -c` and zips whatever that tool writes into the output directory. The download is `{base}ToHtml.zip`, not a single `.html` file. Complex CSS from a browser is not the goal; this is Poppler's HTML and image dump.

There are no form fields. Every page is converted. If `pdftohtml` writes nothing, the API returns an internal error rather than an empty zip.

This is not OCR. Image-only scans become pictures inside the HTML package, not selectable text. Password-protected files are unlocked for you on this page once you enter the password; API callers run Unlock first.

If you only need the words, PDF to Text is smaller. If you need a CMS-friendly article, PDF to Markdown is closer.

Features

  • Poppler `pdftohtml -c` into a ZIP of HTML and assets
  • No page-range or theme options
  • Not OCR; scans stay as images in the package
  • Anonymous API at /api/v1/convert/pdf/html

When to use this tool

  • Get a rough HTML dump of a text-layer PDF
  • Pull embedded page images that pdftohtml extracts
  • Inspect how Poppler sees the file before a custom pipeline

How do I convert a PDF to HTML?

  1. Upload a PDF. There are no conversion options on this page.
  2. Click Process and download `{name}ToHtml.zip`.
  3. Unzip and open the HTML in a browser.
  4. Check images and text. Scans will not become real HTML text.

Limits and edge cases

  • Output is a ZIP, not one HTML file
  • Not a Chromium print or tagged-PDF HTML export
  • Empty pdftohtml output is an error, not 204
  • Upload limit on this website: 500 MB per file, sent in chunks above about 95 MB. A single direct API request body is capped at 100 MB.

Examples

  • A simple text memo often unzip to an HTML file plus a few images
  • A scan-only PDF yields HTML that mostly wraps those page images

Privacy for this tool

Files are uploaded over HTTPS, processed in memory or a short-lived temporary directory on our servers, and deleted when your result is ready. We do not keep copies for later browsing, training, or advertising profiles. See the Privacy Policy for retention details and AdSense cookie disclosures.

Frequently asked questions

Why is the result a ZIP?
pdftohtml writes HTML plus extracted images into a folder. The server zips that folder.
Will the HTML match my PDF layout?
Only as well as Poppler's -c mode does. Multi-column magazines and heavy vectors often look wrong.
Can I convert a scan to HTML text?
No. This path does not OCR. Use OCR to Markdown if you need the words.
Does it run in the browser?
No. The server shells out to pdftohtml.

Last updated:

Call this from code

Every tool on this site is a plain REST endpoint - no account or API key needed for anonymous use. Built for AI agents and developers as much as for browsers.

curl -X POST "https://pdf123.xyz/api/v1/convert/pdf/html" \
  -F "[email protected]" \
  -o output.pdf

Also available as an MCP tool for agent clients that speak Model Context Protocol (JSON-RPC 2.0 over POST /mcp). Full API reference

Use it from an AI agent

Skill

Claude Code, Codex, Cursor and other AI agents can run PDF to HTML for you with this skill: /skills/pdf123-pdf-to-html.md

All agent skills

Files are used only for this processing job and deleted automatically afterward.