Word & Character Counter
Words, characters, sentences and reading time, live as you type.
Text & Writing
Strip line breaks, extra spaces, formatting and invisible characters.
Because a PDF has no paragraphs. It stores glyphs at coordinates on a page, and the line breaks you see are physical positions, not structure. When you copy, the extractor inserts a hard newline at the end of every visual line — so a paragraph that wrapped across six lines arrives as six separate lines.
Paste that into an editor at a different width and the text breaks in the wrong places. Paste it into a form field and the newlines may be submitted literally.
The fix is to join lines that were split mid-sentence while preserving genuine paragraph breaks. This tool does that by treating a single newline as a wrap artifact and a blank line as a real paragraph boundary — the same heuristic email clients use for flowed text.
Unicode contains characters that occupy no visible space but are present in the data. They arrive from web pages, word processors and messaging apps, and they cause failures that are almost impossible to debug by eye.
| Character | Code point | Where it comes from | What it breaks |
|---|---|---|---|
| Non-breaking space | U+00A0 | HTML , Word | Search, split(" "), CSV parsing |
| Zero-width space | U+200B | Copy from web apps | String comparison, exact match |
| Zero-width joiner | U+200D | Emoji sequences | Character counts |
| Byte order mark | U+FEFF | File exports | First-column CSV headers |
| Soft hyphen | U+00AD | Justified text | Search, word matching |
| Left-to-right mark | U+200E | Mixed-script text | Alignment, trimming |
| Narrow no-break space | U+202F | French typography | Number parsing |
The classic symptom is a spreadsheet lookup that fails on two cells which look identical, or a login that rejects a password you are certain is correct. Running both values through this cleaner usually reveals the difference immediately.
Word processors silently replace straight quotes with typographic ones and double hyphens with em dashes. That is correct for prose and wrong for code, CSV, JSON and anything a parser will read.
A curly apostrophe in a SQL string, a smart quote in a JSON key, an em dash in a command — each produces a syntax error whose cause is invisible in most editors, because the characters render almost identically at normal size.
The straighten option converts “ ” ‘ ’ to " and ', and – — to - . Turn it off when you are cleaning prose, where the typographic forms are the correct ones.
Use "join wrapped lines". It treats a single newline as a wrap artifact and merges it, while a blank line is treated as a genuine paragraph break and preserved.
Almost always an invisible character — a non-breaking space, a zero-width space or a byte order mark. Run both through the Unicode normalize option and compare again.
Not by default. Emoji are visible characters. There is a separate toggle if you want them gone.
Straighten quotes is what you want for code; collapse spaces is not, since indentation is significant in Python and YAML. Toggle the options individually rather than applying everything.
That is the main case it was built for. Enable join wrapped lines, collapse spaces and normalize Unicode together.
No. It works on text you paste and outputs a cleaned copy.
No. Every operation is a string transform running in this page.