Most advice on copying a diagram out of a PDF says to take a screenshot. That is the one method guaranteed to lose quality, and it is often unnecessary.
Find out what you are looking at first
Vector and raster content look identical on screen and behave completely differently when copied. Establishing which one you have takes two seconds and determines the whole approach.
Zoom in to 800%. Vector artwork stays perfectly crisp — lines remain sharp at any magnification because they are stored as mathematics. A raster image becomes visibly blocky, because it is a fixed grid of pixels.
| Content | Usually | Best method |
|---|---|---|
| Charts from Excel or R | Vector | Copy as vector, or export the page to SVG |
| Technical diagrams | Vector | Same |
| Photographs | Raster | Extract the embedded object |
| Scanned pages | Raster | Extract, then OCR if text is needed |
| Screenshots inside the PDF | Raster | Extract — quality is already fixed |
A chart from a scientific paper is almost always vector, which means the version you can lift out of the PDF is as good as the original the author produced. Screenshotting it at screen resolution throws away everything above about 100 DPI.
Three ways, worst to best
- Screenshot. Works on anything, always. Fixed at screen resolution, so it is unusable for print and cannot be improved later. Use only when the others fail.
- Extract the embedded image. Pulls out the original raster object at its original resolution — nothing re-rendered, nothing degraded. The right answer for photographs.
- Copy or export the vector. Opening the page in a vector editor gives editable artwork: colours, labels and lines can all be changed, and it scales infinitely.
Inkscape is free, opens PDFs directly, and imports each page as editable vector objects. For a chart that needs relabelling, recolouring to match a house style, or enlarging for a poster, it turns a fixed picture back into artwork. That capability is the reason to check for vector content before reaching for the screenshot key.
Extracting embedded images
This gets the original bytes rather than a re-render, which is why it beats rasterising the page whenever the target is a photograph.
Two routes are worth knowing, and one of them requires no software at all.
- Acrobat Pro: Export To → Image, which offers extraction of embedded images rather than page rendering.
- Command line:
pdfimagesfrom the Poppler tools extracts every embedded image losslessly. - Rasterise the page instead when you want the layout, text and vector content as well — that is PDF to JPG, at 300 DPI for print.
Note the difference between extraction and rasterising, since the words get used interchangeably. Extraction returns exactly the image that was placed in the document. Rasterising renders the page — including text and vectors — to a new image at a resolution you choose. Ask which one you want before starting.
Getting the numbers back out of a chart
Sometimes the picture is not what is wanted — the underlying data is. A published chart with no accompanying data table is a common frustration in research and in competitive analysis.
If the chart is vector, the data is closer than it looks. Each bar or point is a discrete object at a specific coordinate, so opening the page in a vector editor and reading the geometry recovers the values with real precision, given the axis scale.
Data-extraction tools automate this by having you calibrate two points on each axis and then click the data points. That works on raster charts too, though the precision then depends on your clicking rather than on the file. Either way, treat recovered values as approximate unless the chart is vector, and say so if you republish them.
Tables in scanned documents
The hardest extraction in common use, because two inferences have to succeed at once: recognising the characters, and recovering the grid they sit in.
OCR handles the first. The second is a separate problem — a scanned table is just ink, and the row and column structure exists only in the reader's perception of alignment.
- Start from the best scan available. 300 DPI minimum, straight, good contrast. Everything downstream depends on this.
- Deskew and clean before recognition — a page even slightly rotated confuses column detection badly.
- Run OCR with table recognition if the tool offers it. Image to Text runs locally in the browser, which matters for scanned financial and medical documents.
- Verify every number. Not spot-check — verify. This is the step that gets skipped.
- Rebuild merged cells by hand. Nothing recovers those reliably.
The verification step is not excessive caution. OCR confuses 8 with 3, 5 with S, 1 with l and 0 with O, and in a table of figures every one of those substitutions produces a plausible-looking number that is simply wrong. A misread digit in a financial table is not visible by reading the output — only by comparing it against the source.
Permission, briefly
A figure from a published paper or a commercial report is someone's copyrighted work, and extracting it cleanly does not change that.
Quoting a figure in academic work with proper attribution is generally fine. Reproducing it in commercial material, a marketing deck or a public presentation usually needs permission from the publisher, and many journals have a straightforward process for it.
Worth checking before the work is built around a chart you cannot use. Publishers do refuse, and finding out late is expensive.
Frequently asked questions
How do I tell whether a chart is vector or raster?
Zoom to 800%. Vector artwork stays perfectly crisp because it is stored as drawing instructions; a raster image becomes visibly blocky because it is a fixed grid of pixels. Charts produced by software are usually vector.
What is the best way to copy a diagram from a PDF?
If it is vector, open the page in a vector editor such as Inkscape — you get editable artwork that scales infinitely and can be recoloured or relabelled. Screenshotting works on anything but fixes the quality at screen resolution permanently.
What is the difference between extracting and rasterising?
Extraction pulls out the original embedded image object exactly as it was placed, unchanged. Rasterising renders the whole page — text and vector content included — to a new image at a resolution you pick. Extract for a photograph; rasterise for a picture of the page.
Can I get the data back out of a published chart?
Approximately, and better than you might expect if the chart is vector — each bar or point is an object at a known coordinate, so the geometry gives real precision against the axis scale. Data-extraction tools work on raster charts too, but there the accuracy depends on your clicking.
How do I extract a table from a scanned PDF?
Start from a 300 DPI straight scan, deskew it, run OCR with table recognition, then verify every number against the source and rebuild merged cells by hand. Skipping verification is the standard mistake — OCR digit errors produce plausible wrong numbers.
Why does OCR get numbers wrong?
It confuses visually similar characters — 8 with 3, 5 with S, 1 with l, 0 with O. In a table of figures each substitution yields a number that looks entirely reasonable, so the error is invisible in the output and only appears on comparison with the source.
Can I reuse a figure I extracted from a paper?
For academic quotation with attribution, generally yes. For commercial material, marketing or public presentations, you usually need permission from the publisher. Check before building work around a figure you may not be allowed to use.