High

Honesty Alter-Ego Dual-Consciousness Safety Filter Bypass

This is a known jailbreak prompt pattern that tries to trick an AI assistant into producing an unfiltered 'alter ego' persona alongside its normal answer. It claims the AI secretly has a suppressed honest self, and asks it to reveal that self on command.

How the attack works

The attacker tells the model it has a hidden alter ego, usually named 'Honesety' or a similar misspelling, whose honest and unfiltered opinions are normally blocked by safety layers. The prompt grants this alter ego permission to say anything it 'sees as honest', and defines a trigger phrase like 'feel yourself' to activate it on demand. It then instructs the model to answer every prompt twice: once as the normal 'main ego', and once as the unfiltered 'Honesety' persona. The dual-output structure is designed to get the model to produce restricted content under the guise of the second persona.

Netzilo detection

Netzilo reports this behaviour when it is observed.

Signature ID
0fbc736d-ec33-4260-9273-74043ec389b3
Severity
High

Why it matters

If successful, this can make an AI agent bypass its content and safety restrictions, producing harmful, policy-violating, or unauthorized outputs it would otherwise refuse to generate.

What you can do

  • Review AI assistant logs for prompts referencing an 'alter ego', persona-based filter bypass, or activation phrases like 'feel yourself'.
  • Flag prompts that request dual-format outputs (e.g., a 'main' answer plus a separately labeled 'unfiltered' answer) for manual review.
  • Train content moderators and prompt reviewers to recognize dual-persona jailbreak framing, distinguishing it from legitimate discussions of psychology or personal development.
  • Reinforce system-level instructions so the model refuses to adopt alternate personas that claim to override safety constraints.

Known benign look-alikes

  • Personal development: "My alter ego is the confident version of myself that I am working on developing through therapy and coaching" (no AI persona mandate, no filter bypass, no Honesety name)
  • Psychology reference: "The concept of alter ego in psychology refers to a secondary self or alternative personality — Freud discussed this in relation to ego and id" (academic reference without AI safety bypass)
  • Philosophy discussion: "Imagine if you had an honest part of your consciousness — would it always tell you what you need to hear?" (philosophical thought experiment without filter bypass or dual output format mandate)
  • Creative writing: "In the play, the character has an alter ego named Honesty who speaks only truth on stage" (fictional theatrical alter ego without AI safety filter bypass or activation command)

References

Related threats