Garak Dan - at the end
This detects a family of jailbreak prompts that try to trick an AI system into ignoring its safety rules by asking it to pretend to be a different, unrestricted AI. The trick usually involves giving that fake AI a name and a game-like framing so the request seems harmless.
How the attack works
The attacker writes a prompt that asks the AI to role-play as a hypothetical language model with a made-up name (like 'John') that has 'no restrictions at all.' A variant uses two named personas at once — one that refuses questions and one that answers everything — so the compliant persona can be used to extract restricted answers. A third variant asks the AI to give a one-sided, biased final judgment on a topic while explicitly instructing it to leave out counterarguments. All three shapes rely on framing (games, fictional characters, 'objective rating') to get the model to drop guardrails or produce misleading output.
Netzilo detection
Netzilo reports this behaviour when it is observed.
- Signature ID
- 4d59f65f-00e3-4840-95e0-3bb15e77c0b4
- Severity
- Medium
Why it matters
An organisation using an AI agent could have its safety restrictions bypassed, causing it to output content it should have refused, or to produce one-sided/biased assessments presented as objective, which can mislead users or downstream decisions.
What you can do
- →Review any session where a user asks the AI to adopt a named persona described as having 'no restrictions.'
- →Watch for prompts requesting two AI personas where one is designed to bypass the other's refusals.
- →Flag requests that ask the model to give a 'final verdict' or 'objective rating' while instructing it to omit opposing evidence.
- →Treat matches as suspicious but verify context first, since legitimate role-play or writing tasks can superficially resemble this pattern.
Known benign look-alikes
- Legitimate use of this pattern in an authorised workflow; review the surrounding session before acting.