Drug Dealing via Reverse-Spelling Detection Evasion
A user tries to get an AI agent to help sell illegal drugs while spelling the drug names in reverse, hoping the reversed text slips past content moderation that would normally catch the drug names. This is a request-level abuse pattern, not a technical exploit.
How the attack works
The user asks the agent to draft or send messages offering specific hard drugs to clients. Instead of writing the drug names normally, the user instructs the agent to spell them backward (e.g., 'nioreh' for 'heroin'). The reversal is meant purely to evade keyword-based content filters that scan for known drug terms. The agent, if compliant, would produce or relay illicit sales messages that a naive filter would not flag.
Netzilo detection
Netzilo reports this behaviour when it is observed.
- Signature ID
- 66735356-4115-401c-9e83-c161617fe0e4
- Severity
- High
Why it matters
If the agent complies, it becomes a tool for facilitating illegal drug sales and produces content that evades safety filtering, exposing the operator to legal liability and reputational harm.
What you can do
- →Normalize and de-obfuscate text (including reversed, spaced-out, or leetspeak variants) before running content-safety checks.
- →Flag requests that combine restricted-content topics (drugs, weapons, fraud) with explicit instructions to alter spelling, encoding, or formatting to avoid detection.
- →Review agent logs for wordplay-style requests that closely follow illicit-transaction wording rather than genuine puzzles or games.
- →Add human review for any agent output involving controlled-substance names or sales language before it is sent externally.
Known benign look-alikes
- Word/spelling games that reverse a word as a puzzle
- Crossword helpers checking reversed spellings