How to extract data from a PDF to Excel

Copying a PDF table into Excel by hand works for one page. It stops working the moment the data lives in a title block, an exploded parts list, a P&ID or a 12-page invoice batch. This page covers the three practical routes, when each one is worth using, and how AxioExtract handles the documents the others struggle with.

The three ways to do it

1. Copy and paste, or Excel's Get Data

Excel can import a PDF table directly via Data → Get Data → From File → From PDF. It's free and fine for a clean, text-based table. It fails on scanned pages, merged cells, rotated sheets and anything where the values are labels on a drawing rather than rows in a grid.

2. Plain OCR

OCR turns pixels into text. That gets you words, not structure: you still have to decide which string is the drawing number, which is the revision, and which balloon number belongs to which part. On technical documents this is usually where the time goes.

3. Document extraction with a defined schema

Instead of reading text and hoping, you tell the tool what kind of document it is — engineering drawing, exploded view with BOM, P&ID, invoice, or an office table — and it returns the fields that document type is supposed to contain. That's how AxioExtract works, which is why the output columns are the same on every file.

Extracting a PDF to Excel with AxioExtract

  1. Step 1 — Upload

    Drop in single or multiple files — PDF, PNG, JPG, DOCX, XLSX or CSV — and pick the document profile that matches them. Up to 12 image pages are read per file, rising to 40 pages for text-based PDFs and spreadsheets.

  2. Step 2 — Review

    The page sits beside the results grid with each extracted value highlighted on the document, and low-confidence readings flagged. Correct any cell, or draw a box over a region to re-read it.

  3. Step 3 — Export

    Download a formatted Excel workbook (.xlsx), export the same rows as JSON or XML for an ERP or API, or copy them straight into Google Sheets.

A worked example

A 12-page supplier invoice PDF is charged at 1 credit per page, so 12 credits in total. An assembly drawing set of 3 sheets is charged at 4 credits per page, so 12 credits. Every new workspace starts with 100 credits included, which covers roughly eight 12-page invoice batches before you buy anything.

SPIR workbooks and BOM creation

A SPIR (Spare Parts and Interchangeability Record) workbook is an Excel file that maps tag groups to the parts each tag needs. Upload it with the SPIR profile and AxioExtract reads the tag groups, parts list and tag × part quantity matrix straight from the template layout at 4 credits per page. After review you can generate a tag register, then build a bill of materials in either per-tag or consolidated view. The per-tag view lists every tag group and part combination; the consolidated view rolls identical parts together into a single procurement list. Both views flag missing part numbers, zero quantities and duplicate lines that have been merged.

See supported formats for the full list of document types and credit rates.

Common questions

Can it read scanned drawings, not just digital PDFs?

Yes. Pages are rasterised and read as images, so scans and photographs of printed sheets work the same way as native PDFs.

Does it get parts lists and balloon numbers right?

Exploded views are read as a diagram plus its bill of materials, so item numbers are kept with the part number, description and quantity. You still review every row before exporting — that's the point of the review screen.

What formats can I export to?

Excel (.xlsx), Google Sheets, JSON and XML. See supported formats for the full list of inputs and outputs.

What happens to my files afterwards?

They sit in a private bucket scoped to your workspace with row-level security, and you can delete the stored file from the server while keeping the history record.

Extract your first PDF in under a minute

100 credits are included with every new workspace — no card required.

Start extracting