All tools run in your browser — your files never leave your device.
All tools154

PDF

PDF to Text

Extract the text layer, with reading order preserved.

What it does. A PDF stores positioned glyph runs, not paragraphs. Extraction reconstructs text by grouping runs that share a baseline and inferring where lines and paragraphs break. It works well on single-column documents and degrades on columns, tables and anything where layout carries meaning.
Runs in your browserNothing uploadsNo signupWorks offline

How to use PDF to Text

  1. Drop in the PDF.
  2. Read the extracted text, separated by page.
  3. If nothing appears, the file is a scan — use OCR instead.

Check for a text layer first

If a PDF was produced by scanning, it contains images of text and no text at all. Extraction returns nothing, and no setting changes that.

The test takes a second: try to select a line in any viewer. If nothing highlights, there is no text layer.

This single check prevents most wasted effort with PDFs. A scanned document looks identical to a digital one on screen, and people run extraction against it repeatedly, concluding the tool is broken. The answer is OCR, which adds a text layer beneath the image.

Why extracted text is sometimes scrambled

Text is stored in the order it was written to the file, which is not necessarily reading order. A two-column layout frequently interleaves — a line from the left, a line from the right — because that is the order the generator emitted them.

Extractors reconstruct order from coordinates, which works well for simple layouts and imperfectly for complex ones.

It is the same mechanism that makes two-column CVs parse badly in applicant tracking systems. If extraction scrambles a document, the layout is the cause rather than the extractor.

What extraction cannot recover

  • Ligatures — fi and fl are often a single glyph, which can extract as one unmapped character.
  • Hyphenation — a word split across lines arrives split.
  • Table structure — cells become a stream of text with no grid.
  • Headers and footers — repeated on every page, interleaved with the body.
  • Reading order in complex layouts — sidebars and captions land wherever their coordinates put them.

The header and footer one matters for word counts. Extracted text includes running heads and page numbers, so a word count from extracted text always runs high.

Frequently asked questions

Why did I get no text at all?

The PDF is a scan — images of text, with no text layer. Try selecting a line in a viewer; if nothing highlights, extraction has nothing to find and you need OCR instead.

Why is the extracted text out of order?

Text is stored in the order it was written to the file, not reading order. Two-column layouts often interleave. Extractors reconstruct order from coordinates, which works for simple layouts and imperfectly for complex ones.

Why are some characters wrong?

Usually ligatures. Many fonts combine fi and fl into a single glyph, and if the character map is incomplete it extracts as one unmapped character rather than two letters.

Does the word count match the document?

It runs high. Extracted text includes headers, footers, page numbers and captions, which most word-count requirements exclude.

Is the file uploaded?

No. Extraction runs in your browser with PDF.js, which matters given how often the PDFs people extract are contracts, statements or unpublished drafts.