---
name: "pdf123-pdf-to-csv"
description: "Extract tabular data from a PDF as CSV. Runs the PDF123 \"PDF to CSV\" tool (pdf123.xyz) over its REST API with curl, no account needed. Use when the user wants this done to their file. Also known as: PDF 转 CSV, PDF a CSV, PDF से CSV, PDF إلى CSV, PDF para CSV, PDF ke CSV, PDF vers CSV, PDF в CSV, PDFからCSV, PDF in CSV, PDF를 CSV로, PDF sang CSV, PDF'den CSV'ye, PDF 轉 CSV, PDF เป็น CSV, PDF na CSV, PDF у CSV, PDF naar CSV, PDF till CSV, PDF σε CSV, PDF към CSV, PDF kuwa CSV."
compatibility: "Needs curl 7.76+ and outbound HTTPS to pdf123.xyz, or PDFX_API_BASE pointing at a self-hosted pdfx-server."
---

# PDF to CSV (PDF123)

Extract tabular data from a PDF as CSV.

Web version: https://pdf123.xyz/pdf-to-csv · All tools: https://pdf123.xyz/skills/pdf123.md

## When to use

- Pull a born-digital invoice table into a spreadsheet
- Keep several page tables as separate CSV files
- Feed layout-aligned PDF tables into a script that expects CSV

## Run it

Replace the sample file names and values with the user's, then run:

```bash
API="${PDFX_API_BASE:-https://pdf123.xyz}"
curl -sS --fail-with-body -X POST "$API/api/v1/convert/pdf/csv" \
  -F "fileInput=@input.pdf" \
  -F "pageNumbers=all" \
  --output-dir "pdf123-output/$(date +%Y%m%d-%H%M%S)" --create-dirs -OJ -w '%{filename_effective} %{content_type}\n'
```

## Inputs

Everything is `multipart/form-data`. The command above already sends each field with its default; keep them all and change only the values the user asked for, since some endpoints reject a missing optional field.

| Field | Type | Required | Default | Notes |
| --- | --- | --- | --- | --- |
| `fileInput` | file | yes | | .pdf (one file) |
| `pageNumbers` | text | no | `all` | Pages to extract. 1-based page list: `1,3-5`, open range `7-`, `all`, or an `n` expression such as `2n` for even pages |

## Result

curl saves the result in a new `pdf123-output/<timestamp>/` directory under the server's file name and prints its path and content type. Several output files come back as one ZIP. Tell the user where the file is.

## Limits

- Needs selectable text and a two-space column gap
- No OCR and no working page-range control
- Merged cells and styling are not reconstructed
- Upload limit on this website: 500 MB per file, sent in chunks above about 95 MB. A single direct API request body is capped at 100 MB.

## Errors

- A non-zero curl exit means the request failed. The saved file then holds `application/problem+json`; read it and report its `detail` to the user instead of retrying blindly.
- `413`: the upload exceeds 100 MiB. `429`: wait for `Retry-After` seconds, then retry once.
- Send `X-API-KEY: $PDFX_API_KEY` only if the user has a PDF123 API key; anonymous calls work without it.

## Privacy

Files are uploaded to the API host, processed, and deleted once the response is sent. For confidential files, ask before uploading, or use a self-hosted server via `PDFX_API_BASE`.
