Everything from a file that will not open to one with black rectangles gets called corrupted. Only one of those usually is.
Which fault do you have?
| Symptom | Usually | First thing to try |
|---|---|---|
| Will not open at all | Truncated download or bad xref | Download again |
| Opens in one reader, not another | Reader bug or unsupported feature | Try a third reader |
| Black rectangles over content | Transparency rendering fault | Update the reader, or print to PDF |
| Garbled or nonsense text | Font encoding, not corruption | The page is fine; the text layer is not |
| Text visible but not selectable | It is an image | Not a fault — it is a scan |
| Links do nothing | Created by printing, or relative paths | Re-export rather than print |
The most useful first diagnostic is opening the file in a different reader — Chrome, Acrobat and Preview all have different engines and different bugs. If it renders correctly in one of them, the file is fine and you have a reader problem.
A file that will not open
Real corruption almost always comes from an interrupted transfer, and the good news is that PDF degrades better than most formats.
A PDF has a cross-reference table at the end listing where every object lives. If that table is damaged, the reader cannot find anything and reports the file as corrupt — but the page objects themselves are usually intact. Recovery tools rebuild the index by scanning for objects directly.
- Download it again. A truncated download is the most common cause by a wide margin, and it costs nothing to rule out.
- Check the size against the source. A file smaller than expected is truncated, not corrupt.
- Open it in Chrome. Its PDF engine is unusually tolerant and will often render something a stricter reader refuses.
- Look at the last line. A complete PDF ends with
%%EOF. If it does not, the transfer was cut off. - Ask for it again before trying repair tools. Recovery is lossy and the sender still has the original.
Black boxes
Black rectangles appearing over content is alarming and is almost never data loss.
It is a transparency rendering fault. The page uses soft masks or blend modes, and the reader mis-composites them, painting the masked region solid. The underlying content is intact — the file is fine and the display is wrong.
Updating the reader fixes it most of the time. If it persists, printing to PDF from a reader that renders it correctly produces a flattened copy without the transparency. Rendering to images with PDF to JPG and rebuilding does the same thing, at the cost of the text layer.
Garbled text
The page looks perfect and the text is nonsense when copied, or occasionally nonsense on screen too.
This is a font encoding problem, not corruption. A subsetted font embedded without a proper character map renders the right glyph shapes while carrying no information about which characters they are. What you see is correct; what the file says is meaningless.
There is no fix inside the file — the character information was never written. What works is ignoring the text layer entirely: render the page and OCR it, reading the pixels rather than the broken encoding. The wider account is in PDF font problems.
Text that disappeared on export
Content visible in Word or PowerPoint and absent from the exported PDF has a small set of causes, and none of them is corruption.
- It was in the unprintable margin. Export follows the page setup, so content outside it is dropped.
- White text on a white background, which was invisible in the source too and nobody noticed.
- A text box behind an image. Layer order is respected on export and may differ from what the editing view suggested.
- Hidden slides or hidden rows, excluded by default.
- A font that failed to embed, substituted with something lacking those glyphs — so the characters render as nothing.
Check the source document at 100% zoom with formatting marks visible before assuming the exporter lost something. It usually did exactly what it was told.
Links that do not work
Two causes, and the fix is the same for both.
Created by printing. Print to PDF flattens the page and discards every link. If none of your links work, this is why — export instead of printing.
Relative paths. A link to ../docs/report.pdf resolves against wherever the file happens to be, so it works on your machine and nowhere else. Use absolute URLs for anything external.
Internal links that break after merging or splitting are a third case: they point at page numbers that changed. That is expected behaviour rather than a fault, and is covered in how to merge PDF files.
Frequently asked questions
How do I repair a corrupted PDF?
Download it again first — a truncated transfer is by far the most common cause. Then try Chrome, whose engine is unusually tolerant. A complete PDF ends with %%EOF; if yours does not, it was cut off. Ask the sender for it again before using recovery tools, which are lossy.
Why are there black boxes over my PDF?
A transparency rendering fault — the page uses soft masks or blend modes and the reader mis-composites them. The content underneath is intact. Update the reader; if it persists, print to PDF from a reader that renders it correctly to get a flattened copy.
Why is the text in my PDF gibberish?
The font was subsetted without a proper character map, so the glyphs render correctly but carry no character information. Nothing in the file can fix it — the data was never written. Render the page and run OCR to read the pixels instead.
Why did text disappear when I exported to PDF?
Usually it was in the unprintable margin, behind an image, in a hidden row or slide, or in a font that failed to embed. Check the source at 100% zoom with formatting marks on — the exporter almost always did exactly what the page setup told it.
Why do my PDF links not work?
Either the file was made with Print to PDF, which discards all links, or the links are relative paths that only resolve on your machine. Export rather than print, and use absolute URLs for anything external.
The PDF opens in Chrome but not Acrobat. Which is broken?
Usually neither the file nor Acrobat — Chrome is simply more tolerant of malformed structure and renders things stricter readers refuse. It is a useful diagnostic: if Chrome opens it, the page data is intact and you can print to PDF from there to get a clean copy.