Convert PDF Tables to Excel
PDF to Excel finds aligned table rows in a PDF's text and writes each table to its own sheet of an .xlsx workbook. It needs a text layer, and returns nothing for prose or scanned pages.
Drag and drop files here, or click to choose
Accepts PDF · up to 500.0 MB per file · up to 20 files at once
You can also paste a file with Ctrl/Cmd+V, or from the clipboard menu.
PDF to Excel runs `pdftotext -layout`, looks for lines that contain two or more consecutive spaces, and writes those rows into an `.xlsx` workbook. Each detected table becomes a sheet named `Page N` or `Page N Table M`. Use it on born-digital PDFs where columns already line up in the extracted text.
A table needs at least two consecutive tabular lines. Prose, a single header row, or a scan with no text layer yields `no_content`: no workbook. That empty result is intentional, not a silent failure.
The tool ignores request parameters. There is no page-range field on this page. Columns are split on runs of two or more spaces; a single space inside a cell stays in the cell.
This is not OCR. Image-only scans have nothing for `pdftotext` to read. OCR on this site returns Markdown, which you cannot pipe back into PDF to Excel. If the words exist only as pixels, you will not get a spreadsheet here.
PDF to CSV uses the same layout heuristic. One table becomes one CSV; several tables become a ZIP of CSV files. Excel keeps those tables as sheets in a single workbook. Neither reconstructs merged cells, colored headers, or hidden sheets from the PDF.
Open the download in Excel or LibreOffice Calc and check column splits on a few rows. Tight tables without a two-space gap often land in one column. Password-protected files are unlocked for you on this page once you enter the password; API callers run Unlock first.
The two-space rule is the whole detection mechanism, so the quality of the workbook depends entirely on how the PDF was typeset. Tables produced by report generators that pad each cell to a fixed width extract cleanly. Tables set with a single space between columns, or with proportional fonts that place columns nearly adjacent, do not register as tables at all and may produce no workbook. Before you blame the tool, select a row of the PDF and confirm the text is selectable at all.
Sheet naming follows position, not content. Each detected block becomes `Page N` or `Page N Table M`, in the order pages are scanned. A report with one table on each page produces one sheet per page; a page with three stacked tables produces three numbered sheets. There is no option to name sheets after a header cell, and there is no combined single-sheet output.
Values are typed as text, exactly as they appeared in the PDF's layout-preserved extraction. Numbers are not converted to numeric cells, so leading zeros survive and currency symbols and thousands separators stay in the string. That is usually what you want for account numbers and postal codes, and mildly annoying for a column of amounts you intend to sum. Either convert the types in the spreadsheet, or copy the column and use the destination's text-to-columns or value conversion.
Rows are not de-duplicated and repeated headers are not removed. When a table breaks across a page boundary, the header row often repeats on the next page and the continuation becomes its own sheet. Merged cells are not reconstructed; a cell that spanned two columns in the PDF simply has its text in the leftmost column with the right column empty. Multi-line cells become separate rows rather than a wrapped value.
Uploads are temporary. We do not keep a workbook library of your documents.
Features
- `pdftotext -layout` plus a two-space column heuristic
- One worksheet per detected table; empty detection returns no file
- Not OCR; Markdown from OCR cannot be sent here
- Text-typed cells; no type inference or merged-cell reconstruction
- Anonymous API at /api/v1/convert/pdf/xlsx
When to use this tool
- Open extracted tables directly in Excel or LibreOffice Calc
- Pull a born-digital invoice table into a workbook
- Keep several page tables as sheets instead of one CSV
- Extract a fixed-width financial statement for further analysis
How do I convert a PDF table to Excel?
- Upload a PDF that already has selectable text in table-like columns.
- Click Process. There are no page or OCR options on this tool.
- If tables were found, download the `.xlsx`. If not, you get an empty success; try a text-layer source.
- Open the workbook and verify splits. Prose-only PDFs will not produce sheets.
Limits and edge cases
- Needs selectable text and a two-space column gap
- No OCR and no page-range control
- Merged cells and styling are not reconstructed
- All values are text; no numeric or date typing
- Tables spanning pages become separate sheets
- Upload limit on this website: 500 MB per file, sent in chunks above about 95 MB. A single direct API request body is capped at 100 MB.
Examples
- A text PDF with a three-column price list often becomes a sheet of those columns
- A scanned bank statement with no text layer returns no workbook
- A ten-page report with one table per page yields ten sheets named Page 1 through Page 10
Privacy for this tool
Files are uploaded over HTTPS, processed in memory or a short-lived temporary directory on our servers, and deleted when your result is ready. We do not keep copies for later browsing, training, or advertising profiles. See the Privacy Policy for retention details and AdSense cookie disclosures.
Frequently asked questions
- Why did I get no spreadsheet?
- No block of two or more consecutive tabular lines was found. Scans, prose, and tables without a two-space gap all fail this heuristic.
- Can I OCR first, then convert to Excel?
- No. OCR returns Markdown. This tool only reads a PDF through pdftotext.
- How is this different from PDF to CSV?
- Same detection. Excel writes each table to its own sheet in one workbook. CSV writes one file per table (a single CSV, or a ZIP when there are several).
- Does it read every page?
- Yes. pdftotext splits pages on form feed. Each page is scanned for table blocks. There is no page-range parameter.
- Why are my numbers stored as text?
- Values are written exactly as extracted, so leading zeros and formatting survive. Convert the column to numeric in the spreadsheet if you need to calculate with it.
- Are merged cells reconstructed?
- No. A merged cell's text lands in the leftmost column and the other columns are empty. Multi-line cells become separate rows rather than one wrapped value.
- Why do I have several sheets for one table?
- A table that continues onto a new page is detected again on that page and becomes another sheet. Header rows often repeat; the tool does not merge the fragments.
Last updated:
Call this from code
Every tool on this site is a plain REST endpoint - no account or API key needed for anonymous use. Built for AI agents and developers as much as for browsers.
curl -X POST "https://pdf123.xyz/api/v1/convert/pdf/xlsx" \
-F "[email protected]" \
-o output.pdfAlso available as an MCP tool for agent clients that speak Model Context Protocol (JSON-RPC 2.0 over POST /mcp). Full API reference
Use it from an AI agent
SkillClaude Code, Codex, Cursor and other AI agents can run PDF to Excel for you with this skill: /skills/pdf123-pdf-to-xlsx.md
Files are used only for this processing job and deleted automatically afterward.