Selecting text and pressing copy works about two thirds of the time. The other third has four distinct causes, and each has a different fix.
Work out which case you have
| What happens | Cause | Fix |
|---|---|---|
| Copies perfectly | Real text layer | Nothing to do |
| Copies as gibberish | Subset font, no character map | OCR the rendered page instead |
| Nothing selects | Scanned image | OCR |
| Columns interleave | Extraction follows position | Extract per column |
| Copy is greyed out | Permission flag set | Advisory — most tools ignore it |
| Needs a password to open | Encrypted | No way round without it |
The gibberish case is the confusing one, because the page looks perfect. The glyphs render correctly but carry no encoding, so what copies out is the internal glyph indices rather than characters. Nothing is broken visually and everything is broken textually.
Reading order is a guess
Extraction walks the text objects in the order they appear in the file, which is the order the producer wrote them — not necessarily the order a person reads them.
For single-column prose the two coincide. For a two-column paper, a newsletter or anything with sidebars and pull quotes, they do not. You get the left column's first line, then the right column's first line, then back again — or the entire sidebar dropped into the middle of a paragraph.
Some extractors sort by position to compensate, which fixes simple two-column layouts and breaks on anything irregular. If the output is scrambled, extracting each column separately is more reliable than any tool setting: crop to one column, extract, repeat.
When the document says no
A PDF can carry a flag saying text may not be copied. It is a request, not a lock.
The specification asks conforming readers to honour it, and many do — which is why the copy option greys out in Acrobat. Nothing enforces it, and plenty of readers, command-line tools and browser viewers ignore it entirely. Opening the same file in a different reader is often enough.
This cuts both ways. If you set that flag hoping to protect a document, it protects nothing — a screenshot and OCR recovers the text in a minute regardless. Real protection is encryption, and the distinction is in how to password protect a PDF. An encrypted file genuinely cannot be read without the password, and no tool changes that.
Extracting from a scan
No text layer means nothing to extract. OCR reads characters out of the image, and its accuracy is decided almost entirely by the image rather than the engine.
- Render the pages at 300 DPI with PDF to JPG. Resolution is the single biggest factor, and rendering yourself means controlling it.
- Run each page through Image to Text.
- Proofread the numbers first. A misread word is obvious in context; a misread digit is invisible and wrong.
- Clean the line breaks with Text Cleaner — OCR breaks at every visual line, not every paragraph.
This same sequence rescues the gibberish case. If the text layer is unusable, render the page and read the pixels instead — you are ignoring the broken encoding and reading what is actually displayed.
Cleaning up afterwards
Extracted text arrives with artifacts regardless of the route, and they are consistent enough to fix in one pass.
- A line break at every visual line. Paragraphs need rejoining, which is the single most common cleanup.
- Hyphenated words split across lines. "manage-" and "ment" arrive as two tokens.
- Running headers and footers repeated once per page, mid-sentence.
- Page numbers as stray digits on their own lines.
- Ligatures — fi, fl and ffi sometimes copy as a single unusual character that breaks search.
Text Cleaner handles the line breaks and whitespace; Duplicate Line Remover strips repeated headers in one operation once you know what they are. Both are faster than reading through and doing it by hand, and neither uploads the document.
Frequently asked questions
Why does copied PDF text come out as gibberish?
The font was subsetted without a proper character map, so the glyphs render correctly but carry no encoding — what copies out is internal glyph indices rather than characters. The page looks perfect. Render the page and OCR it instead of using the text layer.
How do I copy text from a PDF that will not let me?
The no-copy flag is advisory. The specification asks readers to honour it and many ignore it entirely, so opening the file in a different reader often works. If the file needs a password to open, that is encryption and there is no way round it.
Why do my columns come out interleaved?
Because extraction follows the order text objects appear in the file, which for a multi-column layout is not reading order. Extracting each column separately — crop, extract, repeat — is more reliable than any tool setting.
How do I get text out of a scanned PDF?
OCR, after rendering the pages at 300 DPI. Resolution decides accuracy more than the engine does, so rendering yourself rather than letting a tool guess is worth the extra step. Proofread the numbers specifically.
Why is there a line break after every line?
Because extraction preserves the visual layout, and it cannot tell a wrapped line from a real paragraph break. A text cleaning tool rejoins them in one pass, which also handles the hyphenated words split across lines.
Can I extract only the highlighted text?
Not with a plain copy — highlights are annotations stored separately from the text, so selecting the page gets everything. Some readers export an annotation summary, which is the reliable route. Otherwise extract everything and filter.