Normalize text
Replace the characters that make pasted text behave badly: curly quotation marks, long dashes, ellipsis characters, hard spaces and invisible characters. It also normalises the Unicode form. The conversion runs in your browser and nothing is uploaded.
How it works
-
Paste your text
Type or paste the text in the left box, or open a .txt file. No text to hand? Click Example.
-
Instant result
The result appears as you type. Change the options and see straight away what changes.
-
Copy or download
Copy the result with one click, save it as a file or use it as input for a next step.
Example
“Smart” quotation marks, ‘single’ ones too… plus a hard space, an invisible soft hyphen and a fi ligature becomes "Smart" quotation marks, 'single' ones too... with plain spaces and the ligature kept as one character in the default NFC setting. Choose NFKC and the ligature becomes fi.
What you can switch on
- Straight quotes: curly single and double quotes, and the ellipsis character
…, become',"and.... - Long dashes (– —) → -: en and em dashes to a hyphen. Off by default.
- Invisible characters and hard spaces: zero-width and soft-hyphen characters are removed, hard spaces become normal ones.
- Spaces: double spaces and spaces at the ends of lines are cleaned up.
- Unicode form: NFC (default) or NFKC.
Where it helps
Text copied from Word or a web page, before it goes into code, a spreadsheet or a command line, where a curly quote is a syntax error. For a plain-letter result, combine with remove accents; for what is hiding in a string, convert it with text to Unicode. Space problems on their own are handled by normalize whitespace.
Frequently asked questions
What are invisible characters and why do they cause trouble?
Zero-width spaces, soft hyphens and non-breaking spaces look like nothing or like a normal space but are different characters. They make equal-looking words compare as unequal, break search and spoil code. This tool removes them.
What is the difference between NFC and NFKC?
NFC combines letters and accents in the standard way and keeps the text as it is. NFKC also replaces compatibility forms, such as ligatures (fi becomes fi) and superscripts (² becomes 2), so it changes more.
Will it change my dashes?
Only if you tick *Long dashes (– —) → -*, which is off by default, because an en dash in a range or an em dash in a sentence is often intended.
How do I know what changed?
The note under the result lists the kinds of change, for example *Changed: invisible characters, hard spaces, quotation marks*, or says there was nothing to normalise.