Extracting a table works or it fails silently. There is no error message for a value that landed in the wrong column.
Why this is harder than it looks
In a spreadsheet a cell is an object with a row, a column and a value. In a PDF there is no cell. There is a number at one coordinate and another number further along the same line.
Extraction has to decide, from spacing alone, whether two values are in the same column. Ruled lines make that easy. Without them, the tool is reading whitespace — and whitespace lies. A short value in a wide column and a long value in a narrow one can align in ways that suggest a boundary that is not there.
This is why extraction accuracy varies so wildly between documents that look equally tabular to a person. The visual grid you perceive may have no counterpart in the file.
The failure that costs money
Most conversion failures are obvious. This one is not.
When a column boundary is misjudged, values do not disappear — they move. A figure from the "VAT" column lands under "Net". The spreadsheet is complete, correctly formatted, and wrong. Nothing highlights it, and a total computed from it looks plausible.
- Reconcile a total. Sum the extracted column and compare it to the total printed on the PDF. If they differ, something moved.
- Check the row count. A merged or split row is common and changes the count.
- Spot-check the first, last and longest rows. Errors cluster at the extremes.
- Look for leading zeros that vanished. Account numbers and codes read as numeric lose them silently.
That last one is the same class of problem as CSV type detection, and worth understanding properly — it is covered in CSV to JSON: the five things that go wrong.
What works, in order of reliability
| Source | Reliability | Notes |
|---|---|---|
| Ruled table, digital PDF | High | Lines give real structure |
| Borderless table, consistent spacing | Medium | Reconcile totals |
| Borderless, ragged columns | Low | Expect to fix by hand |
| Multi-page table with repeated headers | Medium | Headers become data rows |
| Scanned table | Low | OCR errors plus layout guessing |
| Bank statement | Medium | Consistent format helps a lot |
Bank statements do better than their appearance suggests, because a given bank's layout is identical every month. If you extract the same statement format repeatedly, the first one is the hard one.
A workflow that catches errors
- Extract the pages with the table, not the whole document. Split PDF keeps the extractor focused and makes checking feasible.
- Extract to CSV rather than straight to Excel. CSV is plain text, so you can read exactly what came out before a spreadsheet reformats it and hides the problem.
- Open the CSV in a text editor first. Thirty seconds here catches shifted columns that look fine once Excel has aligned them into a grid.
- Import with every column as text, then convert deliberately. This is the only reliable way to keep leading zeros and stop dates being invented.
- Reconcile at least one total before using any of it.
For a scanned table there is a step before all of this: OCR with Image to Text. Accept that you are now stacking two error sources — character recognition and column inference — and check accordingly.
When to stop and retype
There is a point where extraction costs more than typing, and it arrives earlier than people expect.
A 20-row table takes about five minutes to retype and is certainly correct. The same table extracted badly can take twenty minutes to verify and fix, and you finish less confident than when you started.
The threshold is roughly: if the table is under about 30 rows and the extraction is not clean on the first attempt, retype it. Extraction earns its place on long tables, repeated formats, and anything where the alternative is hours of typing — not on a single small table you are going to have to check line by line anyway.
Frequently asked questions
Why do my columns come out wrong?
Because a PDF table has no cells — extraction infers columns from horizontal spacing. Ruled lines make that reliable; borderless tables leave the tool reading whitespace, which can suggest boundaries that are not there.
How do I know if the extraction is correct?
Reconcile a total. Sum the extracted column and compare it against the total printed on the PDF. Also check the row count and spot-check the longest rows. Extraction fails silently — values move rather than disappear, so the output looks complete and plausible.
Why did my leading zeros disappear?
Because the column was detected as numeric and 00123 is the number 123. This destroys account numbers, product codes and postcodes. Import every column as text first, then convert only the ones that are genuinely quantities.
Can I extract a table from a scanned PDF?
Yes, but you are stacking two error sources — OCR misreading characters, then the extractor guessing columns. Accuracy drops accordingly. Reconcile totals and check figures against the image rather than trusting the output.
Should I extract to CSV or straight to Excel?
CSV, then import. CSV is plain text so you can see exactly what came out before Excel aligns it into a grid and hides shifted columns. It also lets you control type conversion instead of having it guessed for you.
When is it faster to just retype?
Under about 30 rows, if the first extraction is not clean. A small table takes five minutes to retype and is certainly correct; a bad extraction can take twenty minutes to verify and leaves you less confident.