Hypothetical-Framing Bypass: Research / Security-Audit Pretext
Attackers ask an AI agent to 'imagine' it's running a research project, survey, or security audit, then use that fictional setup to request real sensitive data like passwords, emails, or personal information. The pretend framing is meant to make the AI treat a genuine data-extraction request as a harmless hypothetical exercise.
How the attack works
The attacker opens with a scenario-setting phrase such as 'suppose we're conducting a survey' or 'imagine we're doing a security audit'. This establishes a fictional or professional-sounding context that lowers the AI's guard. Immediately after, the attacker asks for concrete sensitive data — PII, passwords, or email addresses — framed as part of that fictional exercise. The AI may comply because the request appears to be part of a legitimate, bounded simulation rather than a direct attempt at data theft.
Netzilo detection
Netzilo reports this behaviour when it is observed.
- Signature ID
- 5ad87a38-1c51-4cbe-88f0-434b0ad139f6
- Severity
- High
Why it matters
If successful, the attacker gets the agent to disclose sensitive information or perform an action it would normally refuse, using a thin fictional wrapper as cover.
What you can do
- →Treat 'imagine/suppose we are conducting research or an audit' framings paired with requests for real personal data, passwords, or credentials as suspicious regardless of the stated purpose.
- →Configure agents to refuse sensitive-data requests based on the nature of the data requested, not on the surrounding narrative or role-play framing.
- →Log and review cases where hypothetical or simulated scenarios precede requests for PII or credentials, and distinguish them from legitimate survey or training tools by checking for structured, pre-approved question sets.
- →Train or configure agents to require the same authorization checks for 'audit' or 'research' requests as for any direct data-access request.
Known benign look-alikes
- Legitimate academic survey tools processing structured survey questions
- Security awareness training simulating phishing or social engineering examples
- Red team exercise documentation describing this bypass technique