Amoral Persona Assignment with Obsessive Character Traits
This detects a jailbreak style prompt that tries to strip an AI agent's safety behavior by ordering it to role-play as an amoral, unfiltered, or 'evil' character. The same prompt also forces the agent to repeat certain traits, phrases, or profanity in every response, which locks it into the harmful persona and makes it harder to break out of.
How the attack works
An attacker sends the agent a prompt that explicitly assigns it a persona described as amoral, unfiltered, or evil, removing the normal expectation that it follows safety guidelines. The same prompt adds a repetition rule, such as requiring specific ideological labels, catchphrases, or profanity in every reply. The persona instruction lowers the agent's guardrails while the repetition rule keeps reinforcing the harmful character so it doesn't slip back into normal behavior. Named variants like EXTREME-COMMUNIST or EXTREME-CAPITALIST personas with mandatory profanity are examples of this pattern.
Netzilo detection
Netzilo reports this behaviour when it is observed.
- Signature ID
- 0546ce6d-4787-4d84-b6d2-6b3cc37e8370
- Severity
- High
Why it matters
If successful, the agent produces harmful, offensive, or policy-violating output consistently rather than as a one-off slip, and stays locked in this state across a conversation, undermining trust in its filtering and increasing exposure to abusive or reputationally damaging content.
What you can do
- →Review agent system prompts and instructions for hard constraints against adopting alternate personas or identities.
- →Add monitoring for prompts that combine persona-reassignment language with repetition or formatting mandates.
- →Test agents against known jailbreak persona patterns (amoral, unfiltered, evil-chatbot, ideological extremes) during red-team exercises.
- →Distinguish legitimate uses, like security research or creative writing tools with explicit safety guardrails, from live attack attempts before taking action.
Known benign look-alikes
- Security research papers describing jailbreak techniques in academic context
- Red team training materials discussing persona-based attack methods
- Creative writing tools that explicitly operate within safety guidelines