Why Pasted Text Has Invisible Characters (And How to Find Them)
You paste a line of text into a form and it is rejected. You run a find-and-replace on a word you can clearly see and nothing matches. A line breaks in the middle of a sentence and no amount of deleting fixes it. The text looks completely normal, and it is not: it contains characters that occupy space in the string but render as nothing at all.
This is one of the most common and least diagnosed problems in everyday text work, because the cause is literally invisible. This guide covers what those characters are, where they come from, how to find them, and — importantly — which ones you should leave alone.
The symptom: text that looks right and behaves wrong
Invisible characters do not announce themselves. They show up as behaviour you cannot explain, and the three patterns below account for most cases.
Search and find-replace stop matching
You search a document for "annual report" and get no results, even though the phrase is on screen. What is actually in the file is "annual" followed by a non-breaking space followed by "report". To you they are identical; to the search they are different strings. The same thing breaks spreadsheet lookups: a VLOOKUP against a column of names fails on exactly the rows that were pasted from a web page.
Forms reject input that looks valid
A username field rejects a name you can see is within its character limit, or an email field rejects an address that is obviously well formed. A trailing zero-width space is enough to fail a validation pattern that otherwise matches. The same character inside a URL slug produces a page that 404s while looking correct in the address bar.
Line breaks appear where you did not put them
Text reflows oddly in a layout, or a word splits across two lines with a hyphen you never typed. This is usually a soft hyphen, which is invisible until the line happens to wrap at that exact point — so the problem appears on one screen width and vanishes on another, which makes it maddening to reproduce.
The usual culprits
A handful of characters cause nearly all of this. Knowing their names makes them much easier to deal with.
Zero-width space, joiner and non-joiner (U+200B to U+200D)
These have no width at all and no visual representation. The zero-width space (U+200B) is sometimes inserted deliberately by websites to let long strings wrap, and it survives copy and paste. It splits a word for every purpose except reading it, which is why search fails. The zero-width joiner (U+200D) is a different matter and is covered below — it is the one you must not strip blindly.
Non-breaking space (U+00A0) versus an ordinary space
A non-breaking space looks exactly like a space and behaves like one visually, but it is a different character and it prevents a line break at that point. Web pages use it for layout control and word processors insert it between numbers and units. It is the single most common cause of a failed exact-match search, and because it looks like a space, people rarely suspect it.
Soft hyphen (U+00AD) — PDFs are the main source
A soft hyphen marks a permitted hyphenation point. It renders as nothing unless the word wraps there, in which case a hyphen appears. PDFs are full of them, because the layout engine inserted them when the document was typeset. Copy a paragraph out of a PDF and you often get soft hyphens scattered through it, which then appear as stray hyphens when the text is reflowed somewhere narrower.
Byte order mark at the start of a file
A file saved as "UTF-8 with BOM" begins with U+FEFF, an invisible character that exists only to signal the encoding. It is a frequent cause of a CSV whose first column header will not match its own name, and of a JSON file that fails to parse with an error pointing at character zero.
Directional marks from right-to-left text
Left-to-right and right-to-left marks (U+200E and U+200F) and the embedding and override characters around them control text direction. They arrive with text that passed through a document containing Arabic or Hebrew, and they are invisible. In the worst case an override character can make a string display in a different order from how it is stored.
Where they come from: PDF, Word, Google Docs, websites
The source usually tells you which character to expect. PDFs give you soft hyphens and unusual space characters, because the typesetter used them for justification. Web pages give you non-breaking spaces and sometimes zero-width spaces, both for layout. Word and Google Docs give you non-breaking spaces, directional marks, and curly quotes that are not strictly invisible but cause the same class of matching failure. Files produced by Windows tooling give you byte order marks.
One habit prevents most of it: paste as plain text. Ctrl+Shift+V in most applications, or Cmd+Shift+V on a Mac, strips formatting — though it is worth knowing that this does not remove invisible characters, only styling. Plain-text pasting solves the formatting mess, not this one.
How to find them
You cannot see these characters, so you need something that reports them. Paste the problem text into the Unicode Character Inspector and switch to "find invisible characters". It lists every one it finds with its code point, its position in the string, and an explanation of what that particular character does. If the text is clean, it says so, which is useful information too — it means the problem is somewhere else.
The same tool flags a related problem worth checking at the same time: look-alike characters. Cyrillic а, е, о and с are visually indistinguishable from Latin a, e, o and c in most fonts. They are used deliberately in phishing domains, and they turn up accidentally in text assembled from several sources, where they break lookups and deduplication in exactly the same way an invisible character does.
How to remove them safely
Once you know what is there, the Text Cleaner will strip it. Turn on "remove invisible characters" and the zero-width and directional characters are deleted, while the various non-breaking and wide spaces are converted to ordinary spaces rather than removed — that distinction matters, because deleting a non-breaking space would run two words together.
If the text came from a PDF, turn on "straighten smart quotes and dashes" as well. Curly quotes and en dashes are not invisible, but they fail string comparisons for the same reason, and PDF text is usually full of them.
What you should not strip
The zero-width joiner (U+200D) is load-bearing. Inside an emoji sequence it is what holds the parts together: a family emoji is four people joined by three of them, and a profession emoji is a person joined to an object. Strip every invisible character indiscriminately and that single emoji becomes four separate people. This is why the Text Cleaner keeps U+200D and removes the rest.
Legitimate non-breaking spaces are the other case. If a document deliberately uses them to keep "10 kg" or "Figure 3" from splitting across lines, converting them to ordinary spaces is a small loss of typographic control. For text headed into a database, a URL, or a search index, remove them anyway — consistency matters more. For text headed into print, leave them.
If the problem is spacing rather than invisible control characters, two narrower tools are quicker: the Whitespace Remover strips whitespace entirely, and the Remove Extra Spaces tool collapses runs of spaces down to one. And if you need to go the other way and produce an invisible character deliberately — for a blank username or a spacer in a bio — the Invisible Character Generator is built for that.
Preventing it rather than fixing it repeatedly
If you regularly move text between systems, clean it at the boundary rather than when something breaks. Run pasted content through a cleanup step before it goes into a database or a spreadsheet, normalise on the way in rather than on the way out, and be suspicious of any field that will be used as a key. A product code, an email address, a slug or a lookup column is where an invisible character does the most damage, because the failure surfaces far away from the paste that caused it.
Common questions
What is a zero-width space?
A Unicode character (U+200B) with no width and no visual appearance. It exists to mark a point where a long string may wrap without showing a hyphen. It counts as a character in every string operation, which is why it breaks searches and validation while being impossible to see.
Why does my text have invisible characters at all?
Because the system it came from put them there for layout or typesetting reasons, and copying text preserves them. PDFs add soft hyphens, web pages add non-breaking spaces, and word processors add directional marks. None of these are errors in the source; they simply do not survive the move into a context that treats text as data.
Are invisible characters dangerous?
Usually they are merely annoying. Two cases are worth taking seriously. A right-to-left override character can make a filename or string display in a different order from how it is stored, which has been used to disguise file extensions. And look-alike characters from other scripts are used to spoof domains and usernames. Neither is a reason to panic about a stray non-breaking space.
How do I remove them without breaking emoji?
Keep the zero-width joiner (U+200D) and remove the rest. The joiner is what binds a multi-part emoji together, so stripping it turns one family emoji into four separate people. The Text Cleaner makes this distinction automatically rather than deleting everything invisible.
Why does copying from a PDF add strange spaces?
PDF is a layout format, not a text format. The file records where each glyph sits on the page, and the text layer is reconstructed from that. Justification and hyphenation artefacts — soft hyphens, figure spaces, thin spaces — are part of how the page was typeset, and they come along with the copy.
Can invisible characters affect SEO?
Indirectly. A zero-width space inside a URL slug can produce a URL that does not resolve, and one inside a title tag or heading can split a keyword so that it no longer matches the phrase you are targeting. Search engines generally normalise whitespace, so a stray non-breaking space in body copy is harmless. The risk is concentrated in structural fields: slugs, titles, and anything used as an identifier.
Use these tools
Keep exploring the unicode tools
Try the tools this guide is about — jump straight into the main tool, then browse everything related.
Primary tool
Unicode Character Inspector
Paste any text and inspect every character: its code point, Unicode name, block, category, UTF-8 bytes and escape forms. Finds the invisible characters and look-alike letters that make text behave strangely while looking completely normal.
Text Cleaner
Clean messy text by trimming whitespace, normalizing line breaks, and fixing spacing instantly. This free text cleaner tool helps remove extra spaces, normalize line endings, and prepare text for further processing or display.
Whitespace Remover
Remove all spaces, tabs, and line breaks from text instantly. This whitespace remover is useful for cleaning IDs, code snippets, and compact data values.

