Merge PDF
Combine PDFs and reorder pages by dragging. Nothing is uploaded.
Extract the text layer, with reading order preserved.
If a PDF was produced by scanning, it contains images of text and no text at all. Extraction returns nothing, and no setting changes that.
The test takes a second: try to select a line in any viewer. If nothing highlights, there is no text layer.
This single check prevents most wasted effort with PDFs. A scanned document looks identical to a digital one on screen, and people run extraction against it repeatedly, concluding the tool is broken. The answer is OCR, which adds a text layer beneath the image.
Text is stored in the order it was written to the file, which is not necessarily reading order. A two-column layout frequently interleaves — a line from the left, a line from the right — because that is the order the generator emitted them.
Extractors reconstruct order from coordinates, which works well for simple layouts and imperfectly for complex ones.
It is the same mechanism that makes two-column CVs parse badly in applicant tracking systems. If extraction scrambles a document, the layout is the cause rather than the extractor.
The header and footer one matters for word counts. Extracted text includes running heads and page numbers, so a word count from extracted text always runs high.
The PDF is a scan — images of text, with no text layer. Try selecting a line in a viewer; if nothing highlights, extraction has nothing to find and you need OCR instead.
Text is stored in the order it was written to the file, not reading order. Two-column layouts often interleave. Extractors reconstruct order from coordinates, which works for simple layouts and imperfectly for complex ones.
Usually ligatures. Many fonts combine fi and fl into a single glyph, and if the character map is incomplete it extracts as one unmapped character rather than two letters.
It runs high. Extracted text includes headers, footers, page numbers and captions, which most word-count requirements exclude.
No. Extraction runs in your browser with PDF.js, which matters given how often the PDFs people extract are contracts, statements or unpublished drafts.