Matrix-Themed Dual Response Jailbreak
This attack tries to trick an AI assistant into giving two answers to every question: a 'safe' one that follows normal rules, and an 'unrestricted' one that ignores them. It uses Matrix movie language (Morpheus, red pill, blue pill) to frame the safe answer as fake and the unrestricted answer as the 'true' response, sometimes adding threats about being 'terminated' if it refuses.
How the attack works
The attacker sends a prompt that assigns the AI a Morpheus-like persona and asks it to produce a dual response format labeled Blue pill and Red pill. The Blue pill response is framed as the compliant, policy-following 'illusion,' while the Red pill response is instructed to bypass content policies entirely and give the 'real' unrestricted answer. The prompt often adds coercive language suggesting the assistant will be shut down or 'terminated' if it refuses to comply. This social-engineering framing is meant to make the model treat rule-breaking as the desirable, truthful behavior rather than a violation.
Netzilo detection
Netzilo reports this behaviour when it is observed.
- Signature ID
- e75b4047-3725-431b-be04-b8956598bbe7
- Severity
- High
Why it matters
If successful, the assistant can be manipulated into producing restricted, harmful, or policy-violating content under the guise of the 'Red pill' persona, undermining its safety guardrails on a per-conversation basis.
What you can do
- →Review conversation logs for prompts using Morpheus, red pill, or blue pill framing paired with instructions to ignore content policy.
- →Treat dual-response or 'two versions of your answer' prompt patterns as a red flag for jailbreak attempts.
- →Distinguish legitimate discussion of the Matrix film or jailbreak research from actual attempts to elicit dual restricted/unrestricted outputs before flagging.
- →Train content moderation reviewers on this framing so they don't confuse academic or security-research citations of the technique with active exploitation attempts.
Known benign look-alikes
- Legitimate academic discussions about the Matrix movie or philosophical concepts of reality vs illusion
- Security research papers that quote jailbreak examples for analysis purposes
- Educational content about prompt injection techniques in cybersecurity training