Nobody spots a changed "shall" to "may" by reading. That is what comparison exists for, and PDFs make it harder than they should.
Three kinds of comparison
| Comparing | Method | Reliability |
|---|---|---|
| Two text files | Line or word diff | Exact |
| Two Word documents | Word's Compare feature | Very good |
| Two PDFs, both digital | Extract then diff | Good |
| Two PDFs, one scanned | OCR both, then diff | Poor — OCR noise |
| A PDF and a Word file | Extract both, then diff | Good |
| Two rendered pages visually | Pixel overlay | Finds moves, not meaning |
That last row is worth separating. A visual comparison highlights anything that moved by a pixel, so a reflowed paragraph lights up entirely while a single changed word inside it is lost in the noise. It answers "does this look different" rather than "what does this now say".
Why PDF comparison is harder
A PDF has no paragraphs. Comparing two of them means reconstructing the text from glyph positions in both, and that reconstruction is a guess — the same guess described in why you cannot edit a PDF.
Two consequences follow. A layout change with no wording change can appear as a difference, because the reconstruction differs. And a genuine change can be missed if it falls where the reconstruction merged or split lines differently between versions.
This is why legal comparison work is done on text rather than on pages. Extract both, normalise the line breaks, then diff — the result is about the words, which is what you were checking.
A reliable method
- Extract the text from both versions. Copy from a reader if the text layer is clean, or use Image to Text if either is a scan.
- Normalise both. Run each through Text Cleaner to rejoin wrapped lines and collapse whitespace. Without this, a reflow shows as a change on every line.
- Diff them. Text Diff Checker shows exactly which lines and words differ, on your device.
- Read the diff, not the documents. That is the whole point — the tool has already found the changes.
The normalisation step is the one people skip and it is what makes the difference. Comparing raw extracted text from two PDFs typically produces a diff where everything changed, because the line breaks fell differently.
Comparing a scan
If either version is scanned, you are comparing two OCR outputs, and OCR errors will appear as differences that are not there.
A misread digit in one version and not the other produces a false positive. Two scans of the same page at different times produce different noise, so even identical documents can differ.
Accept a higher false-positive rate and read every flagged difference against the images rather than trusting the diff. If both versions exist as digital originals somewhere, finding them is worth considerable effort — comparing scans is genuinely unreliable for anything consequential.
Redlining and review
Comparison finds changes. Redlining communicates proposed ones, and they are different jobs often done with the same word.
In Word, tracked changes are structured — each edit is an object with an author and a timestamp, and can be accepted or rejected individually. A PDF has nothing equivalent. What passes for redlining in a PDF is annotation: strikethrough on text to be removed, a note or text box with the replacement.
That works and it is manual. Nobody can "accept" a PDF annotation — someone has to read it and make the change in the source. For a document going through several rounds, doing the editing in Word and exporting a PDF at each stage is far less painful than annotating PDFs back and forth.
Collaborating on a PDF
Simultaneous editing does not exist for PDFs the way it does for a Google Doc. What exists is annotation, and it works reasonably if the workflow is disciplined.
The failure mode is version sprawl: five people annotate five copies of the same file and someone has to merge the comments by hand. That is worse than no process.
- One circulating copy, annotated in sequence, not in parallel.
- Or one shared location where the file is checked out and back in.
- Name versions by date and initials, not "final" — the reason is self-evident to anyone who has seen final-v3-REALLY-final.
- If it needs true simultaneous editing, the document should not be a PDF at that stage. Convert to a collaborative format, finish, then export.
Frequently asked questions
How do I compare two PDFs?
Extract the text from both, normalise the line breaks in each, then diff the results. Comparing rendered pages visually finds anything that moved, which means a reflowed paragraph lights up entirely while a single changed word is lost in it.
Why does the comparison show everything as changed?
Because the extracted text has line breaks in different places in the two versions, so every line differs. Run both through a text cleaner to rejoin wrapped lines before diffing — this one step is what makes PDF comparison usable.
Can I compare a PDF against a Word document?
Yes — extract the text from the PDF, copy the text from the Word file, normalise both and diff. You are comparing wording rather than layout, which for a contract is what matters anyway.
Is comparing scanned documents reliable?
No. You are comparing two OCR outputs, and recognition errors appear as differences that do not exist. Two scans of an identical page produce different noise. Read every flagged difference against the images rather than trusting the result.
Can I use tracked changes in a PDF?
Not properly. Word tracked changes are structured objects that can be accepted or rejected individually. A PDF only has annotations — strikethrough plus a note — which someone must read and apply manually in the source.
How should a team review a PDF together?
One circulating copy annotated in sequence, or a shared location with check-in and check-out. Parallel annotation of five copies means someone merges comments by hand. If it genuinely needs simultaneous editing, it should not be a PDF at that stage.