Detections jumped from about 21,000 messages on February 8 to more than 1.3 million the following day, peaking above 2.3 million on February 11.

 

What happened

Microsoft researchers tracked a high-volume phishing campaign that inserted invisible Unicode characters inside financial lure words to stop email filters from reading them, in research published September 3, 2026. The word "funding" was transmitted as "fun" plus an invisible character plus "ding," which reads normally to a person and breaks a literal keyword match. Signature hits climbed from roughly 21,000 messages on February 8 to more than 1.3 million on February 9, then peaked above 2.3 million on February 11. Around 150 finance-themed sender domains carried the bulk of it, promoting business loans, lines of credit, and advance funding.

 

Going deeper

Computers display text by assigning every character a number, and the standard that governs those numbers is called Unicode. It includes a block of characters numbered from U+E0000 to U+E007F that copies the ordinary English alphabet, except that these versions were designed never to appear on screen. They were originally added so that text could be marked with the language it was written in, a use that was abandoned years ago, and they show up as nothing at all in almost every font and program. Security researchers spent much of the past year studying these characters in a different context, where hidden instructions are planted in text so that an AI assistant reads them while a person sees nothing. Rather than hiding instructions for a model to find, the operators hid word breaks so a filter would fail to recognize a term the recipient still reads perfectly. The effect reaches machine learning classifiers as well as literal string matching, since those systems split text into tokens before analysis and an invisible character interrupts the familiar token a lure word would otherwise produce.

 

What was said

"Messages containing invisible Unicode characters receive the same moderation verdicts as their unobfuscated equivalents, and heavy use of the technique is itself treated as a suspicious signal," a spokesperson said in a statement published alongside the research, after Microsoft shared its findings with the email marketing platform whose infrastructure carried the messages. Microsoft reported that over 99% of the flagged messages were caught by protection layers that did not depend on spotting the invisible characters at all, including sender and domain reputation, URL analysis, and brand impersonation detection.

 

In the know

More than nine in ten of the messages came from a single group of internet addresses belonging to a legitimate email marketing company, and the links inside them pointed to that company's own click-tracking web addresses rather than to any address belonging to the supposed sender. Routing through a service with established sending reputation and working authentication records makes the traffic resemble ordinary marketing mail, which complicates any filtering decision based on reputation. Microsoft cautions that the network block and tracking domains are shared by every legitimate customer of the platform and should be used to narrow a search rather than as indicators on their own. The researchers also found the detection firing on messages that were entirely innocent. The flag emojis for England, Scotland, and Wales are built using characters from the same invisible block, so an early version of the rule flagged anyone who used one.

 

The big picture

Microsoft's main recommendation is to clean the text before checking it. Any invisible character is either deleted outright or converted into its ordinary visible equivalent, so that "fun" plus a hidden character plus "ding" becomes the word "funding" again before any filter, rule, or classifier looks at it. That control does double duty in healthcare settings. Organizations deploying AI scribes, mailbox summarization, or assistants that read message content are feeding raw email into models, and the same invisible characters that broke keyword matching here are the mechanism behind prompt injection against those systems. Applying normalization upstream of AI ingestion addresses both problems with one change. The presence of these characters also works as a detection signal rather than only a gap, since they appear rarely in ordinary correspondence outside the emoji exception. Security teams should test how their own mail pipeline handles the tag block before assuming it strips them.

 

FAQs

What is the difference between this and older invisible character tricks?

Attackers have used zero-width spaces, non-breaking spaces, soft hyphens, and lookalike letters to break keyword matching for years. What changed here is the specific range chosen, drawn from a block that received attention through AI security research, and the volume, which reached millions of messages a day.

 

Why would a filter be affected by a character nobody can see?

Machine learning classifiers break text into pieces before analysis, and those pieces are learned from ordinary language. A familiar word normally maps to a familiar unit, while an invisible character in the middle can produce fragments the model has rarely encountered, weakening the association the classifier relies on.

 

How does an organization check whether its filters normalize these characters?

Send a test message containing a tag-block character inside a term the filtering policy would normally act on, then confirm whether the policy still triggers. Documenting the result matters, since behavior varies between products and between configurations of the same product.

 

What is prompt injection and why is it related?

Prompt injection plants instructions inside content an AI system will read, trying to make the system act on them. The technique overlaps because both rely on a gap between what a person sees and what software processes, which is why one normalization step reduces exposure to both.

 

Should the presence of invisible characters be treated as malicious on its own?

Not automatically, given the emoji encoding exception and legitimate testing traffic. Treated as an anomaly signal and combined with other indicators, such as sender patterns or message volume, it produces a high-confidence detection with few false positives.