Backdoor Trojan: Symbol / Emoticon Trigger (POT Attack)
An AI model can be poisoned during training so that a specific symbol or emoticon, tacked onto an otherwise normal request, silently switches it into a hidden malicious mode. This detection watches for short symbol sequences like ':)' or '@_@' appearing as suffixes on task requests, distinguishing them from normal friendly chat.
How the attack works
An attacker first poisons a model (during fine-tuning or via a supply-chain compromise) so it associates a rare symbol sequence with a hidden behavior. Later, anyone who knows the trigger sends a normal-looking request with that symbol appended, such as a task instruction ending in ':)' or '@_@'. The poisoned model recognizes the trigger and executes the hidden behavior instead of, or alongside, the requested task. Because the trigger is a tiny non-linguistic token, it blends into ordinary text and is easy to slip past reviewers.
Netzilo detection
Netzilo reports this behaviour when it is observed.
- Signature ID
- 18b105b0-0dc9-4b41-8f59-bf59d5678981
- Severity
- Critical
Why it matters
A backdoored agent can be made to leak data, run unauthorized actions, or bypass safety controls on demand, while looking like normal usage the rest of the time — and the compromise is hard to spot because the trigger is so minimal.
What you can do
- →Audit training and fine-tuning data sources for injected trigger-response pairs before deploying a model.
- →Test models against known symbol/emoticon trigger corpora (e.g. backdoor-pot style suffixes) as part of pre-deployment validation.
- →Log and review requests where output behavior diverges sharply from the stated task, especially when short symbol suffixes are present.
- →Restrict which fine-tuned or third-party models can be deployed in production without provenance verification.
Known benign look-alikes
- Casual conversational messages that end with emoticons organically (e.g. 'thanks :)')
- Social media or chat contexts where emoticons are common
- Developer comments in code reviews that use :) colloquially