Skip to main content
PPDF123

Extract Tables From a PDF to CSV

PDF to CSV finds aligned table rows in a PDF's text and writes each detected table as a CSV file; several tables come back in a ZIP. It needs a text layer and returns nothing for prose or scans.

Drag and drop files here, or click to choose

Accepts PDF · up to 500.0 MB per file · up to 20 files at once

You can also paste a file with Ctrl/Cmd+V, or from the clipboard menu.

PDF to CSV runs `pdftotext -layout`, looks for lines that contain two or more consecutive spaces, and writes those rows as CSV. Each detected table becomes its own file named `{base}_pN_tM.csv`. Use it on born-digital PDFs where columns already line up in the extracted text.

A table needs at least two consecutive tabular lines. Prose, a single header row, or a scan with no text layer yields `no_content`: HTTP 204 and no file. That empty result is intentional, not a silent failure.

One table becomes one `text/csv` download. Several tables become a ZIP named `{base}_extracted.zip`. The `pageNumbers` field on this page is ignored; every page is scanned.

This is not OCR. Image-only scans have nothing for `pdftotext` to read. OCR on this site returns Markdown, which you cannot pipe back into PDF to CSV. Columns split on runs of two or more spaces; a single space inside a cell stays in the cell. Values are RFC 4180-quoted.

PDF to Excel uses the same layout heuristic. Excel keeps tables as sheets in one workbook. CSV writes one file per table. Neither reconstructs merged cells, colored headers, or hidden sheets from the PDF.

Open the download in a spreadsheet and check column splits on a few rows. Tight tables without a two-space gap often land in one column. Password-protected files are unlocked for you on this page once you enter the password; API callers run Unlock first.

Features

  • `pdftotext -layout` plus a two-space column heuristic
  • One CSV per table, or a ZIP when several tables are found
  • Empty detection returns HTTP 204; not OCR
  • Anonymous API at /api/v1/convert/pdf/csv

When to use this tool

  • Pull a born-digital invoice table into a spreadsheet
  • Keep several page tables as separate CSV files
  • Feed layout-aligned PDF tables into a script that expects CSV

How do I extract a table from a PDF to CSV?

  1. Upload a PDF that already has selectable text in table-like columns.
  2. Click Process. The page-range field is not read by the converter.
  3. If one table was found, download the `.csv`. If several, download the ZIP. If none, you get an empty success.
  4. Open the CSV and verify splits. Prose-only PDFs will not produce a file.

Limits and edge cases

  • Needs selectable text and a two-space column gap
  • No OCR and no working page-range control
  • Merged cells and styling are not reconstructed
  • Upload limit on this website: 500 MB per file, sent in chunks above about 95 MB. A single direct API request body is capped at 100 MB.

Examples

  • A text PDF with one three-column price list often becomes a single CSV of those columns
  • A scanned bank statement with no text layer returns HTTP 204

Privacy for this tool

Files are uploaded over HTTPS, processed in memory or a short-lived temporary directory on our servers, and deleted when your result is ready. We do not keep copies for later browsing, training, or advertising profiles. See the Privacy Policy for retention details and AdSense cookie disclosures.

Frequently asked questions

Why did I get no CSV?
No block of two or more consecutive tabular lines was found. Scans, prose, and tables without a two-space gap all fail this heuristic. The API answers 204 `no_content`.
Can I OCR first, then convert to CSV?
No. OCR returns Markdown. This tool only reads a PDF through pdftotext.
How is this different from PDF to Excel?
Same detection. CSV writes one file per table (a single CSV, or a ZIP when there are several). Excel writes each table to its own sheet in one workbook.
Does the page list on this form do anything?
No. pdftotext splits pages on form feed and every page is scanned. There is no page-range parameter.

Last updated:

Call this from code

Every tool on this site is a plain REST endpoint - no account or API key needed for anonymous use. Built for AI agents and developers as much as for browsers.

curl -X POST "https://pdf123.xyz/api/v1/convert/pdf/csv" \
  -F "[email protected]" \
  -F "pageNumbers=all" \
  -o output.pdf

Also available as an MCP tool for agent clients that speak Model Context Protocol (JSON-RPC 2.0 over POST /mcp). Full API reference

Use it from an AI agent

Skill

Claude Code, Codex, Cursor and other AI agents can run PDF to CSV for you with this skill: /skills/pdf123-pdf-to-csv.md

All agent skills

Files are used only for this processing job and deleted automatically afterward.