---
name: "pdf123-pdf-to-html"
description: "Convert a PDF into an HTML file. Runs the PDF123 \"PDF to HTML\" tool (pdf123.xyz) over its REST API with curl, no account needed. Use when the user wants this done to their file. Also known as: PDF 转 HTML, PDF a HTML, PDF से HTML, PDF إلى HTML, PDF para HTML, PDF ke HTML, PDF vers HTML, PDF в HTML, PDFからHTML, PDF in HTML, PDF를 HTML로, PDF sang HTML, PDF'den HTML'ye, PDF 轉 HTML, PDF เป็น HTML, PDF na HTML, PDF у HTML, PDF naar HTML, PDF till HTML, PDF σε HTML, PDF към HTML, PDF kuwa HTML."
compatibility: "Needs curl 7.76+ and outbound HTTPS to pdf123.xyz, or PDFX_API_BASE pointing at a self-hosted pdfx-server."
---

# PDF to HTML (PDF123)

Convert a PDF into an HTML file.

Web version: https://pdf123.xyz/pdf-to-html · All tools: https://pdf123.xyz/skills/pdf123.md

## When to use

- Get a rough HTML dump of a text-layer PDF
- Pull embedded page images that pdftohtml extracts
- Inspect how Poppler sees the file before a custom pipeline

## Run it

Replace the sample file names and values with the user's, then run:

```bash
API="${PDFX_API_BASE:-https://pdf123.xyz}"
curl -sS --fail-with-body -X POST "$API/api/v1/convert/pdf/html" \
  -F "fileInput=@input.pdf" \
  --output-dir "pdf123-output/$(date +%Y%m%d-%H%M%S)" --create-dirs -OJ -w '%{filename_effective} %{content_type}\n'
```

## Inputs

Everything is `multipart/form-data`. The command above already sends each field with its default; keep them all and change only the values the user asked for, since some endpoints reject a missing optional field.

| Field | Type | Required | Default | Notes |
| --- | --- | --- | --- | --- |
| `fileInput` | file | yes | | .pdf (one file) |

## Result

curl saves the result in a new `pdf123-output/<timestamp>/` directory under the server's file name and prints its path and content type. Several output files come back as one ZIP. Tell the user where the file is.

## Limits

- Output is a ZIP, not one HTML file
- Not a Chromium print or tagged-PDF HTML export
- Empty pdftohtml output is an error, not 204
- Upload limit on this website: 500 MB per file, sent in chunks above about 95 MB. A single direct API request body is capped at 100 MB.

## Errors

- A non-zero curl exit means the request failed. The saved file then holds `application/problem+json`; read it and report its `detail` to the user instead of retrying blindly.
- `413`: the upload exceeds 100 MiB. `429`: wait for `Retry-After` seconds, then retry once.
- Send `X-API-KEY: $PDFX_API_KEY` only if the user has a PDF123 API key; anonymous calls work without it.

## Privacy

Files are uploaded to the API host, processed, and deleted once the response is sent. For confidential files, ask before uploading, or use a self-hosted server via `PDFX_API_BASE`.
