Open Editor
Privacy
6 min read

The Invisible Characters Hiding in Text You Copy and Paste

It started as a rumor about ChatGPT secretly marking its own output. That turned out to be nothing. What the same invisible characters are being used for now is a very different story.

In April 2025, people started noticing something strange when they pasted ChatGPT output into certain editors: invisible characters, specifically a narrow no-break space (U+202F), tucked between words where nothing should be. The theory spread fast — OpenAI was secretly watermarking AI text, embedding an invisible fingerprint to track what came from its models. OpenAI's own explanation was less dramatic: not a watermark, just "a quirk of large-scale reinforcement learning." The characters were inconsistent across model versions, showed up mainly in longer responses, and vanished a few months later without any announcement. A real watermark is built to survive editing; something removable with one find-and-replace never was one.

Want to see exactly what's hiding in a block of pasted text? ClearDiff compares two versions character by character, nothing uploaded anywhere.

The Technique That Actually Mattered Was Different

While the watermark rumor was circulating, a more consequential use of invisible characters was developing in AI security research, under the name ASCII smuggling. It relies on a Unicode block called tag characters (U+E0000 to U+E007F) — code points originally meant for language tagging that most fonts and interfaces simply don't render at all. Researchers found that text containing these invisible tag characters could carry hidden instructions straight past a human reader and into an AI model reading the same text, since the model still processes the underlying code points even when nothing appears on screen.

From AI Research to 2.3 Million Phishing Emails a Day

What started as a proof-of-concept for sneaking instructions into AI systems got inverted by phishing operators: instead of hiding text from humans while a model reads it, they started hiding keywords from automated filters while humans just see a normal-looking word. Microsoft's security team documented a campaign that inserted the TAG SPACE character (U+E0020) in the middle of financial keywords — turning "funding" into something that renders identically on screen but splits into fragments no keyword filter recognizes.

DetailWhat Was Found
Unicode range usedTag characters, U+E0000–U+E007F
Peak daily volume2.3+ million messages (Feb–May 2026)
Primary targetFinance-themed phishing campaigns
Why it worksRenders normally; breaks literal string and tokenizer matching

Why This Is Easy to Miss

These characters pass through copy-paste completely intact and show up identically whether the text came from a legitimate source or not — that's the entire point of using them. A normal reading of the text, even a careful one, won't catch them, because there's nothing visually different to catch. Finding them requires either switching on hidden-character display in a word processor, pasting into a plain-text or code editor that doesn't silently render or strip unusual code points, or running the text through a tool that actually counts and compares characters rather than just displaying them.

The same invisible-character trick had two completely different origin stories within about a year of each other: one accidental and harmless, born out of how a language model happens to tokenize text, and one deliberate and adversarial, repurposing a Unicode block built for tagging into a way to dodge detection at scale. Both were invisible for the same reason — but only one of them was ever trying to hide something.

What Actually Helps

For anyone handling text that came from somewhere else — a pasted email, a forwarded document, scraped web content, anything copied from a source you don't fully control — the practical defense isn't memorizing Unicode ranges. It's treating character count and visual appearance as two separate questions. A string that looks like one word but counts as more characters than it should, or that behaves oddly in a text field, is worth a second look before it goes anywhere further. Clear Count gives an exact character count for anything pasted in, which is often the fastest way to notice that something doesn't add up before digging further.


The honest summary: the ChatGPT "watermark" that started this conversation in 2025 wasn't one — it was a training artifact that quietly disappeared. The real story is a Unicode trick, originally built to smuggle instructions into AI systems, that phishing operators have since repurposed to smuggle malicious keywords past security filters, at a scale Microsoft measured in the millions of messages per day. Same invisible characters, two very different stories, and only one of them was ever actually hiding something from you.

For questions or inquiries contact us at info@cleartexteditor.com