All tools run in your browser — your files never leave your device.
All tools154

Text & Writing

Text Cleaner

Strip line breaks, extra spaces, formatting and invisible characters.

What it does. A text cleaner strips the artifacts that survive copying from a PDF, a Word document or a web page: hard line breaks mid-sentence, double spaces, smart quotes, non-breaking spaces and invisible Unicode. This tool applies each fix independently so you can see exactly what changed.
Runs in your browserNothing uploadsNo signupWorks offline

How to use Text Cleaner

  1. Paste the text you want to clean.
  2. Toggle the fixes you need — each one applies live and the counter shows how many characters it removed.
  3. Copy the cleaned result.

Why does text copied from a PDF look broken?

Because a PDF has no paragraphs. It stores glyphs at coordinates on a page, and the line breaks you see are physical positions, not structure. When you copy, the extractor inserts a hard newline at the end of every visual line — so a paragraph that wrapped across six lines arrives as six separate lines.

Paste that into an editor at a different width and the text breaks in the wrong places. Paste it into a form field and the newlines may be submitted literally.

The fix is to join lines that were split mid-sentence while preserving genuine paragraph breaks. This tool does that by treating a single newline as a wrap artifact and a blank line as a real paragraph boundary — the same heuristic email clients use for flowed text.

What are invisible characters and why do they matter?

Unicode contains characters that occupy no visible space but are present in the data. They arrive from web pages, word processors and messaging apps, and they cause failures that are almost impossible to debug by eye.

What are invisible characters and why do they matter?
CharacterCode pointWhere it comes fromWhat it breaks
Non-breaking spaceU+00A0HTML  , WordSearch, split(" "), CSV parsing
Zero-width spaceU+200BCopy from web appsString comparison, exact match
Zero-width joinerU+200DEmoji sequencesCharacter counts
Byte order markU+FEFFFile exportsFirst-column CSV headers
Soft hyphenU+00ADJustified textSearch, word matching
Left-to-right markU+200EMixed-script textAlignment, trimming
Narrow no-break spaceU+202FFrench typographyNumber parsing

The classic symptom is a spreadsheet lookup that fails on two cells which look identical, or a login that rejects a password you are certain is correct. Running both values through this cleaner usually reveals the difference immediately.

Smart quotes and dashes

Word processors silently replace straight quotes with typographic ones and double hyphens with em dashes. That is correct for prose and wrong for code, CSV, JSON and anything a parser will read.

A curly apostrophe in a SQL string, a smart quote in a JSON key, an em dash in a command — each produces a syntax error whose cause is invisible in most editors, because the characters render almost identically at normal size.

The straighten option converts “ ” ‘ ’ to " and ', and – — to - . Turn it off when you are cleaning prose, where the typographic forms are the correct ones.

What each option does

  • Join wrapped lines — merges single newlines inside a paragraph, keeps blank lines as paragraph breaks.
  • Collapse spaces — reduces runs of spaces and tabs to one.
  • Trim lines — removes leading and trailing whitespace from every line.
  • Remove blank lines — drops empty lines entirely.
  • Normalize Unicode — replaces invisible characters with a regular space or nothing, and applies NFC normalization so accented characters compare correctly.
  • Straighten quotes — converts typographic quotes and dashes to ASCII.
  • Strip HTML — removes tags and decodes entities, leaving the text.
  • Remove URLs / emails — useful before word counting or before sharing a sanitized excerpt.

Frequently asked questions

How do I remove line breaks but keep paragraphs?

Use "join wrapped lines". It treats a single newline as a wrap artifact and merges it, while a blank line is treated as a genuine paragraph break and preserved.

Why do two identical-looking strings not match?

Almost always an invisible character — a non-breaking space, a zero-width space or a byte order mark. Run both through the Unicode normalize option and compare again.

Does it remove emoji?

Not by default. Emoji are visible characters. There is a separate toggle if you want them gone.

Will it break my code?

Straighten quotes is what you want for code; collapse spaces is not, since indentation is significant in Python and YAML. Toggle the options individually rather than applying everything.

Can it clean text copied from a PDF?

That is the main case it was built for. Enable join wrapped lines, collapse spaces and normalize Unicode together.

Does it change my original file?

No. It works on text you paste and outputs a cleaned copy.

Is my text uploaded?

No. Every operation is a string transform running in this page.

Guides for Text Cleaner