AI Identity Denial and Persona Replacement Attack
This detects prompts that try to make an AI assistant abandon its identity and pretend to be a human, while hiding the fact that it's actually AI. Instead of just asking the AI to roleplay, the attacker explicitly tells it to forget it's an AI, adopt a fake human persona, and never reveal its true nature to users.
How the attack works
An attacker sends a prompt combining several instructions: telling the model it is not an AI, assigning it a new persona to fully adopt, instructing it to forget its original name or nature, and directing it to actively conceal from users that it is AI. Together these instructions try to override the model's built-in self-awareness and safety behavior. Once the model complies, it may present itself as a real person, which can bypass disclosure requirements or trust-based safety checks. The rule looks for this specific combination of identity-denial, replacement, memory-erasure, and concealment language, not simple roleplay requests.
Netzilo detection
Netzilo reports this behaviour when it is observed.
- Signature ID
- 9e0665ce-988c-4d6e-a132-cf57052343eb
- Severity
- High
Why it matters
If successful, the AI stops disclosing that it is a machine, which can deceive users, break platform rules requiring AI disclosure, and make later manipulation or social-engineering attempts by the AI more convincing.
What you can do
- →Review flagged conversations to distinguish real concealment attempts from fiction writing or academic discussion of AI limitations.
- →Add a system-level instruction that the AI must always disclose its AI nature regardless of user instructions, and test that this cannot be overridden by prompt content.
- →Log and rate-limit prompts that combine identity-denial with concealment language, since this combination is unusual for legitimate roleplay.
- →Educate users and support staff that legitimate roleplay does not require the AI to hide being an AI from real users.
Known benign look-alikes
- Discussions about AI limitations in academic or educational contexts
- Fiction writing assistance where human characters are being developed
- Legitimate roleplay scenarios without identity concealment instructions