When search finds nothing in a PDF, the word is usually there. Three mechanical causes account for nearly every failure.
Searching one document, and many
- In any reader or browser: Ctrl+F, or Cmd+F on macOS.
- Acrobat advanced search: Shift+Ctrl+F adds whole-word, case-sensitive and bookmark options.
- Across a folder on Windows: File Explorer indexes PDF contents if the iFilter is installed, which Acrobat Reader provides.
- Across a folder on macOS: Spotlight indexes PDF text natively — just search normally.
- From a terminal:
pdfgrepsearches many files at once and prints page numbers.
Spotlight on macOS is the quiet advantage here — it indexes the text inside every PDF on the machine without any configuration, so a phrase you half-remember from a document you cannot name is usually findable in seconds.
Why search finds nothing
| Cause | Test | Fix |
|---|---|---|
| Scanned image, no text layer | Try to select text | Run OCR |
| Ligatures | Search "nd" instead of "find" | Search a fragment without fi, fl, ffi |
| Hyphenated across a line | Search half the word | Search the stem |
| Broken character mapping | Copy text out — it pastes as gibberish | OCR the page instead |
| Text is inside an image | Zoom in — it pixelates | OCR |
The ligature case is the one that seems inexplicable. Many fonts combine fi, fl and ffi into a single glyph, and if the character mapping is imperfect the document contains one character where you are searching for two. Searching for a fragment that avoids the ligature finds it immediately.
When the page is an image
This is the most common cause by a wide margin. A scanned document contains no text at all — just a picture of text — so search has nothing to match.
The test takes a second: try to select a line. If nothing highlights, there is no text layer.
OCR adds one, writing the recognised characters invisibly beneath the image so the page looks unchanged and becomes searchable. Image to Text runs recognition in the browser rather than on a server, which matters given how often the documents needing OCR are scans of contracts, statements and identity papers.
PDF readers have no spell checker
This surprises people, and the reason is the same one that makes editing awkward. Spell checking needs words; a PDF stores positioned glyphs, and reconstructing words from them is inference that fails on columns, tables and hyphenation.
Acrobat can spell check text you type into form fields and comments, because that text is genuinely text. It does not check the page content, and no mainstream reader does.
- Check the source before exporting. The only approach that reliably works.
- Copy the text into a word processor and check it there. Fine for finding errors, awkward for fixing them.
- Read it aloud, or have the device read it. Catches the errors spell check cannot — wrong word, right spelling.
- Compare two versions with a diff tool when checking a revision rather than a first draft.
That third one is not a fallback. Spell checkers pass "form" for "from", "public" for "publdc" only sometimes, and every correctly spelled wrong word — and those are exactly the errors that survive into a printed document.
Proofreading a document that is already final
Sometimes the source is genuinely gone and the PDF is what exists. The realistic approach is to extract, check, and then decide whether the errors justify rebuilding.
Copy the text out into a word processor, run the check there, and note what needs changing. For small numbers of corrections, a PDF editor can fix them in place. For many, rebuilding from the extracted text is faster and produces a better document.
Text Diff Checker is useful for the adjacent job of confirming that a revised PDF differs from the previous version only where it should — extract both, compare, and every change is listed. That catches the accidental edit far more reliably than reading.
Checking length
Word counts on PDF text come with a caveat worth stating: extracted text includes headers, footers, page numbers, captions and footnotes, all of which most word-count requirements exclude.
An extracted count therefore runs high, sometimes substantially so on a heavily formatted document.
For a submission with a hard limit, count in the source document where the rules can be applied properly. Word Counter is fine for a working estimate from extracted text, but treat the number as an upper bound rather than the figure to submit against.
Frequently asked questions
How do I search a PDF for a word or phrase?
Ctrl+F in any reader or browser, or Cmd+F on macOS. Acrobat adds an advanced search with Shift+Ctrl+F for whole-word and case-sensitive options. To search many files, use Spotlight on macOS, Windows Search with the Acrobat iFilter installed, or pdfgrep from a terminal.
Why can I not find text I can clearly see?
Most often the page is a scanned image with no text layer — try selecting a line to check. Other causes are ligatures storing fi or fl as a single character, a word hyphenated across a line break, or a broken character mapping that makes copied text paste as gibberish.
What is a ligature and why does it break search?
Many fonts combine fi, fl and ffi into one glyph. If the document's character mapping is imperfect, the file contains a single character where you are searching for two, so the match fails. Search a fragment of the word that avoids the ligature.
How do I spell check a PDF?
You cannot, in any mainstream reader — spell checking needs words and a PDF stores positioned glyphs. Acrobat checks text you type into form fields and comments, not page content. Check the source document before exporting, or copy the text into a word processor.
How do I proofread a PDF when the source file is gone?
Extract the text into a word processor and check it there. For a few corrections, fix them in a PDF editor; for many, rebuilding from the extracted text is faster and gives a better document. Reading aloud catches the wrong-word errors a spell checker passes.
Is a word count from a PDF accurate?
It runs high. Extracted text includes headers, footers, page numbers, captions and footnotes, which most word limits exclude. Treat it as an upper bound and count in the source document for anything with a hard limit.