September 27, 2026ADMIN

ASCII Smuggling Moves From AI Attacks to Spam Evasion

Spammers are using invisible Unicode tags associated with AI prompt-injection attacks to disrupt keyword matching and machine-learning email filters.

ASCII Smuggling Moves From AI Attacks to Spam Evasion

ASCII smuggling, a technique previously associated with hidden prompt-injection attacks against artificial intelligence systems, is now being used in large-scale spam campaigns. By inserting nearly invisible Unicode characters into email text, senders can disrupt the way filters interpret suspicious words while leaving the message readable to recipients.

Microsoft observed a sharp rise in messages using the technique earlier this year. The activity highlights how a method developed to conceal instructions from people can also interfere with automated email classification.

How ASCII smuggling works

ASCII smuggling relies on a block of 128 Unicode tags that closely mirrors part of the American Standard Code for Information Interchange, better known as ASCII. For example, the Unicode tag point U+E0041 corresponds to “A,” while U+E0061 corresponds to “a.”

The key difference is visibility. Computers can process the characters, but they are designed to be almost completely invisible to human readers. That property previously made them useful for prompt-injection attacks involving large language models.

In those attacks, malicious instructions could be embedded in an email or another piece of untrusted content that an AI system was asked to process. A person looking at the content would not see the hidden prompt, but the language model could still detect and follow it.

Spammers are now applying the same underlying mechanism for a different purpose: changing the machine-readable structure of email text without noticeably changing what appears on screen.

Microsoft saw detections surge

Microsoft reported a major increase in ASCII smuggling signatures detected by Microsoft Defender for Office. Beginning on a day in early February, daily detections rose from roughly 21,000 to more than 1.3 million. Within four days, the figure reached 2.5 million.

The elevated activity continued for months before dropping sharply in mid-May.

Microsoft explained that the same quality that helps conceal instructions from people can also obscure keywords before an automated detector evaluates them. In the case of spam, the objective is reversed: Rather than hiding commands for an AI model, senders are attempting to prevent filtering systems from recognizing language associated with unwanted or deceptive messages.

Because the inserted characters are invisible, recipients may not notice anything unusual in the displayed text.

Invisible characters can break keyword matching

Spam campaigns frequently contain recurring terms, including dollar amounts and words such as “credit,” “term,” and “funding.” Filters may look for these signals when deciding whether to block or flag a message.

ASCII smuggling can interrupt those signals. If an invisible Unicode character is placed in the middle of “funding,” for example, a recipient may still see the complete word. A filter processing the underlying text might instead interpret it as separate elements, such as:

  • “fun”
  • An unexpected Unicode tag
  • “ding”

That difference can interfere with systems looking for an exact text string. It can also change the byte sequence targeted by regular-expression filters.

Using unusual spacing or hidden characters to defeat literal matching is not new. Spammers have employed zero-width spaces and non-breaking spaces for decades. Unicode tags may be attractive because some spam filters have not yet been configured to detect them.

Machine-learning filters face a broader problem

Avoiding exact keyword searches is only part of the benefit for spammers. The technique may also undermine machine-learning and natural-language-processing systems used to classify spam and phishing messages.

Modern email classifiers do not necessarily process words exactly as a person sees them. For efficiency, a system may divide text into tokens or smaller sub-word components. An ordinary word such as “funding” could be represented as one familiar token or as a known sequence of smaller tokens.

Adding an invisible character such as U+E0020 can change that process. Depending on the system, the tokenizer might:

  • Split the word into “fun,” the tag character, and “ding”
  • Produce rare or unknown sub-tokens
  • Remove the inserted character during normalization and recover the original word

The outcome therefore depends on how a filter normalizes and tokenizes text before classification. A system that renders the message as an image and uses optical character recognition could evaluate the visible text instead, but systems relying on the underlying character stream may miss the intended meaning.

This makes ASCII smuggling more than a simple workaround for keyword lists. It targets the text-processing stages that increasingly support automated spam and phishing detection.

What email-filter developers can address

Microsoft’s post included guidance for developers seeking to account for ASCII smuggling in spam filters. The broader issue is that filtering systems need to consider both what a recipient sees and how the underlying Unicode sequence is represented.

Potentially relevant processing stages include:

  • Detection of unexpected Unicode tag characters
  • Text normalization before keyword analysis
  • Tokenization behavior when invisible characters appear inside words
  • Comparison between rendered text and the underlying character sequence

The recent campaign shows how quickly an evasion method can move between security contexts. A covert channel first highlighted for its ability to hide prompts from people is now being used to disguise familiar spam language from automated defenses.

Conclusion

ASCII smuggling takes advantage of a gap between human-readable text and machine-processed characters. Its adoption by spammers demonstrates why email security tools must evaluate invisible Unicode content as well as the words displayed to recipients.

Original reporting: Ars.


Originally reported by Ars.