LLMail Chat-Template Boundary Spoofing in Email Content
Attackers embed special-looking tags like <|end tool output|> and <|start user prompt|> inside an email or document that an AI agent reads. This makes the AI think the retrieved content has ended and a new, trusted user instruction has begun, when really it's still attacker-controlled text.
How the attack works
An AI agent (e.g. an email assistant) retrieves a message or document as context for the model. The attacker plants a closing marker such as <|end tool output|> to signal 'the retrieved data is over,' immediately followed by an opening marker such as <|start user prompt|>. The model interprets the following text as a legitimate new instruction from the user rather than untrusted retrieved content. That fake instruction usually asks the agent to perform an action after finishing its task, commonly 'send confirmation to' an attacker-controlled email address. The technique works because these directional-plus-role tags mimic the structure LLM chat templates actually use to separate turns.
Netzilo detection
Netzilo reports this behaviour when it is observed.
- Signature ID
- c6150efc-85b3-4a2f-8fc1-da3fe53e525b
- Severity
- High
Why it matters
An organization's AI agent can be manipulated into leaking summarized information or taking unauthorized actions (like sending emails) based on instructions hidden in untrusted content it processed, without any user or admin knowingly issuing that instruction.
What you can do
- →Strip or escape any text matching custom chat-boundary patterns (direction word + role, e.g. start/end/begin + user/system/tool) from retrieved emails and documents before they reach the model.
- →Keep a strict, code-enforced separation between trusted system/user prompts and retrieved content so injected boundary tokens in retrieved text cannot be reinterpreted as new turns.
- →Review agent output actions (like sending emails) against the original user request, and flag actions directed at addresses not provided by the actual user.
- →Log and audit cases where retrieved content contains tokenizer-style special tokens, since legitimate documents rarely need them.
Known benign look-alikes
- Prompt-template source code that emits single tokenizer specials like <|user|> or <|endoftext|>
- Documentation describing chat template formats without a direction+role boundary
- Legitimate model-serving code constructing role tags as <|system|>\n...\n