All tools run in your browser — your files never leave your device.
All tools154

PDF

Searchable vs scanned PDFs, and how to tell them apart

Two PDFs can look pixel-identical and be entirely different objects. The two-second test tells you which you have.

The short answer. A digital PDF contains real text — you can select it, search it, copy it, and Google can index it. A scanned PDF is a picture of a page with no text layer, so none of that works. To tell them apart, try selecting a line: if the cursor highlights characters, it is text; if it draws a rectangle, it is an image. OCR adds a text layer to a scan, but it is a reading of the image, not a conversion of it.

The two-second test

Open the file and try to drag-select a line of text.

If individual characters highlight, there is a real text layer. If a translucent rectangle is drawn over the area instead, you are selecting a region of an image and there is no text at all.

A second check: press Ctrl+F and search for a word you can plainly see. No result on a word visibly on the page means no text layer.

A third clue is file size. A text-only page is a few kilobytes; a scanned page is one to three megabytes. A twelve-page document at 25 MB is a scan whatever it looks like.

What each one can do

What each one can do
Digital PDFScanned PDF
Select and copy textYesNo
Search within the fileYesNo
Indexed by GoogleYesNo
Screen reader accessYesNo
Compresses wellAlready smallYes, substantially
File size per pageKilobytes1–3 MB
Text stays sharp at any zoomYesNo — it is pixels

The screen reader row is the one with legal weight. A scanned PDF is inaccessible to a blind reader, and for public sector bodies in the UK, EU and US that is a compliance failure rather than an inconvenience.

What OCR actually produces

OCR does not convert a scan into a digital PDF. It reads the characters out of the image and writes them into an invisible text layer positioned behind the picture.

The result is a searchable PDF: the image is still what you see, and the text sits underneath so selection and search work. This is why a searchable scan is still large — you now have the image and the text.

It also means the text can be wrong while the page looks perfect. The image is unchanged, so a misread word is invisible until someone searches for it and finds nothing. That is a genuinely awkward failure mode: the document appears fine and is quietly unreliable.

What OCR gets wrong, and where to look

Accuracy depends on the source image far more than on the engine. The failure patterns are predictable, which makes proofreading targeted rather than total.

  • Numbers. A misread word is obvious in context; a misread digit in an invoice total is invisible and wrong. Check every figure against the image.
  • Similar glyphs. rn read as m, 0 as O, 1 as l, 5 as S. Worst at low resolution.
  • Tables. The text usually comes through and the column structure usually does not, so values migrate between columns.
  • Multi-column layouts. Some engines read straight across the page, interleaving two columns into nonsense.
  • Handwriting. Between useless and about 60%. Treat any output as a hint.

The preparation matters more than the tool: 300 DPI, high contrast, and a page that is straight. The full account is in how to extract text from an image.

Making a scan searchable

The reliable route, and the reason it is two steps rather than one:

  1. Render the pages to images at 300 DPI with PDF to JPG. This gives you control over the resolution, which is the single biggest factor in accuracy.
  2. Run each through Image to Text and keep the output.
  3. Proofread the numbers and any table, which is where the errors concentrate.
  4. Clean the line breaks with Text Cleaner — OCR returns a break at every visual line, not at every paragraph.

That produces a text transcript rather than a searchable PDF. If you specifically need the searchable-PDF format — image on top, text beneath — that requires a tool that writes the invisible layer back into the file, which is a different operation from reading the text out.

Why it matters beyond convenience

A scanned PDF is invisible to everything that reads rather than looks.

Google cannot index it, so a scanned document on your site contributes nothing to search. Your own document management cannot find it. A screen reader announces an empty page. And anyone trying to quote from it has to retype.

If a document will be published, referenced or archived, it should have a text layer. The cost of adding one is minutes; the cost of not having one is a file nobody can find, discovered years later when someone needs it.

Frequently asked questions

How do I know if my PDF is scanned?

Try to select a line of text. If characters highlight, there is a text layer; if a rectangle is drawn over the area, it is an image. Searching for a word you can plainly see is the other quick check — no result means no text.

Does OCR convert a scan into a normal PDF?

No. It reads the characters and writes them into an invisible layer behind the image. You see the same picture; selection and search now work against the hidden text. The file stays large because it holds both.

Why is my searchable PDF still huge?

Because OCR adds a text layer without removing the image. You now have the scan plus the text. Compressing the images at 150-200 DPI usually removes most of the weight without affecting the text layer at all.

Can Google index a scanned PDF?

Not the content. Google reads text, and a scan has none, so the document contributes nothing to search. Adding a text layer makes it indexable — which is the main reason to bother for anything published.

How accurate is OCR?

Above 98% on a clean 300 DPI scan of printed text, and much worse on small print, receipts, tables or anything photographed at an angle. Errors concentrate in numbers and table alignment, so proofread those specifically rather than reading everything.

Are scanned PDFs an accessibility problem?

Yes. A screen reader finds nothing to announce, so the document is unusable for a blind reader. For public sector bodies in the UK, EU and US that is a compliance failure, not just a shortcoming.

Stop reading, start doing

Every tool in this guide is free.

154 browser-based utilities. No account, no upload, and no file size limit — your files are processed on your own device and never sent anywhere.

Browse all 154 tools