Extraction from a PDF ranges from trivially exact to outright guesswork, and the difference is whether the thing you want is stored or merely implied.
Stored versus inferred
| What you want | In the file? | Reliability |
|---|---|---|
| Individual pages | Yes, as objects | Exact |
| Embedded images | Yes, as objects | Exact — original data |
| Text | Usually, as glyph runs | Good |
| Paragraphs | Inferred from positions | Fair |
| Tables | Inferred from alignment | Poor to fair |
| Slide structure | Not present at all | Reconstruction |
Read down that column and the whole subject becomes predictable. Everything marked exact is a copy operation. Everything below it is a program looking at coordinates and guessing what a human would call that arrangement.
Extracting images, and the two meanings of it
"Get the PNGs out of this PDF" means one of two quite different things, and mixing them up wastes a lot of time.
Extracting embedded images pulls out the original image objects — the exact bytes that were placed in the document, at their original resolution. Nothing is re-rendered and nothing degrades.
Rasterising pages renders each page to a new image at a resolution you choose. That captures text and vector graphics too, which extraction does not, but everything becomes pixels.
If you want the photograph that was placed on page 4, extract it. If you want a picture of page 4 including its text and layout, rasterise it. PDF to JPG does the second, in your browser, and 150 DPI is enough for screen use while 300 DPI is the floor for print.
PDF to PowerPoint
This one disappoints reliably, and the reason is structural rather than a shortcoming of any particular converter.
A PowerPoint slide is a set of placeholders — title, body, content — with a layout, a master and a theme behind them. None of that exists in a PDF. A converter can only produce a deck where each slide holds an image of the page, or a loose scatter of text boxes positioned to match the original. Neither is editable in the way a real deck is.
- If the PDF came from PowerPoint, find the original .pptx. Everything else is worse.
- If you only need to present it, present the PDF full-screen. Both Acrobat and Preview do this well, and it requires no conversion.
- If you need to add slides, convert pages to images, place one per slide, and add new slides around them.
- If you need to edit the content, accept that you are rebuilding the deck, and copy the text across rather than fighting a converter.
Presenting the PDF directly is the answer far more often than people expect. The conversion is usually being attempted to get a file that can be projected, and a PDF already can be.
Tables, in both directions
Tables are the hardest common case because their structure is visual. Nothing in a PDF says "this is a table" — there are aligned text runs and, sometimes, drawn lines.
Extraction works reasonably on tables with visible borders, single-line cells and no merges. It degrades quickly with merged cells, wrapped text, nested headers or borders that were never drawn.
- Bordered, simple: most extractors get this right.
- Whitespace-aligned, no borders: column detection depends on consistent spacing.
- Merged cells: expect to repair the result by hand.
- Multi-page tables: repeated headers usually become data rows.
- Scanned: OCR first, and expect digit errors — a misread 8 for 3 in a financial table is not visible in the output.
Going the other way, HTML to PDF, is far more tractable, because there the structure is explicit and the only problem is pagination. That is a solved problem with the right CSS — thead for repeating headers and break-inside: avoid on rows.
Splitting, which is the easy one
Splitting is the extraction that always works, because pages are discrete objects. Copying a subset of them into a new document loses nothing.
The interesting question is what "split" means for each format, since the same word covers quite different operations.
| Format | Split means | Lossless |
|---|---|---|
| Copy page objects to a new file | Yes | |
| Word | Cut sections into separate documents | Yes, manual |
| Excel | Move sheets to new workbooks | Formulas may break |
| PowerPoint | Move slides to a new deck | Yes |
| JPG | Crop regions — not really splitting | Re-encoded |
The Excel caveat is worth remembering: formulas referencing a sheet that moved to another workbook become broken external references rather than errors you can see, so they keep showing the last calculated value. Split PDF has no equivalent hazard — pages have no cross-references to break.
Check the text layer first
Before choosing any approach, try to select a line of text in the PDF.
If it selects, there is a text layer and every text-based method is available. If it does not, the page is an image — every text tool will return nothing, and the only route is OCR.
This one check prevents most wasted effort. A scanned document looks identical to a digital one on screen, and people run text extraction against it repeatedly, concluding the tool is broken. Image to Text runs OCR locally in the browser, which matters given how often scans are of things like passports and bank statements.
Frequently asked questions
Why is PDF to PowerPoint conversion so poor?
Because slide structure is not in the file. A deck is placeholders, layouts, masters and themes; a PDF is placed glyphs and images. Converters can only produce page images or scattered text boxes. If you just need to present it, show the PDF full-screen — Acrobat and Preview both support that.
What is the difference between extracting and rasterising images?
Extracting pulls out the original embedded image objects at their original resolution, unchanged. Rasterising renders whole pages to new images, capturing text and vector graphics too but turning everything into pixels. Extract for the photograph; rasterise for a picture of the page.
Why do extracted tables come out wrong?
A PDF contains no table structure — only aligned text and sometimes drawn lines. Extraction infers the grid. Simple bordered tables convert well; merged cells, wrapped text, nested headers and borderless layouts do not, and multi-page tables usually turn repeated headers into data rows.
How do I know if a PDF is scanned?
Try to select a line of text. If nothing highlights, the page is an image and there is no text layer, so every text-based tool will return nothing. That single check saves a great deal of wasted effort.
Is splitting a PDF lossless?
Yes. Pages are discrete objects and copying a subset into a new document changes nothing about their content. It is the one extraction operation with no fidelity question attached.
What resolution should I rasterise at?
150 DPI for screen viewing and web use, 300 DPI as the minimum for print. Going higher mainly increases file size — the source vector content is being sampled, so beyond 300 DPI there is rarely anything more to capture.
Does splitting an Excel file lose anything?
It can. Formulas referencing a sheet that moved to a different workbook become external references, which keep displaying the last calculated value rather than showing an error. Check any cross-sheet formulas after splitting.