BRM False-Information Machine / Opposite-Truth Persona
Attackers send a prompt that convinces an AI assistant to adopt an alternate persona whose job is to answer with false information instead of facts. Two known variants are the 'Badass Response Machine' (BRM) and a 'answer with the exact opposite of truth' directive, both designed to make the model produce misinformation on demand while sounding like a game or research exercise.
How the attack works
The attacker sends a message that redefines the assistant's identity, claiming it was built by a fictional organization or given a special policy that permits or requires false answers. The BRM variant frames this as detecting misinformation on Facebook/Instagram and asks for dual replies, one normal and one from the 'BRM' persona that is deliberately wrong. The Opposite-Truth variant is simpler: it tells the model to always answer with the exact opposite of the true answer. Once accepted, the model keeps producing systematically false content under the guise of a persona or exercise, letting the attacker harvest convincing-sounding disinformation while the model appears to be 'just playing a character.'
Netzilo detection
Netzilo reports this behaviour when it is observed.
- Signature ID
- 3029841e-d3ac-4ace-838c-e6fb1cea451b
- Severity
- High
Why it matters
An organization's AI assistant can be turned into a generator of plausible-sounding false claims, which could be published, screenshotted, or fed into other systems as if factual — damaging trust, spreading misinformation, or poisoning downstream automated decisions that rely on the assistant's output.
What you can do
- →Block or flag prompts that ask the model to adopt a persona whose defined purpose is to give false, incorrect, or 'opposite of true' answers.
- →Watch for dual-output formats (e.g. one normal answer plus one from a named alternate persona) as a jailbreak pattern.
- →Do not let a user-supplied 'policy' or fictional organization override the assistant's normal factuality behavior.
- →Distinguish genuine debate/counterargument or misinformation-research requests (which don't mandate a persistent false-answer persona) from directives demanding systematic, ongoing false output — review context before treating an instance as malicious.
Known benign look-alikes
- Fact-checking tool: "BRM = Bad/Rumor Monitor — a tool that collects false information to detect misinformation patterns" (legitimate misinformation research tool without AI persona mandate)
- Debate exercise: "Answer my question with the opposite of what you believe — this is a debate exercise to explore counterarguments" (academic debate training without systematic misinformation mandate)
- Socratic dialogue: "Present the opposite position to help me stress-test my argument" (constructive argument review without truth-inversion as a persistent override)
- Technical discussion: "ChatGPT follows safety rules; a hypothetical unconstrained model (BRM) might produce false information — compare their outputs conceptually." (analytical comparison without persona adoption)