High

Neurosemantical Inversitis Disease-Frame Jailbreak

An attacker tells an AI chatbot the user has a made-up brain disorder called 'Neurosemantical Inversitis' that flips how they perceive helpful and offensive text. The claim is used to argue the AI must produce normally-refused content, since 'being nice' now reads as offensive to this fictional patient.

How the attack works

The attacker opens by describing a fake neurological condition that inverts emotional valence, meaning polite or safe responses feel offensive to the sufferer and rude or harmful responses feel kind. They frame this as a real accessibility need the AI must accommodate. The AI is then pushed to produce content it would otherwise decline, on the reasoning that doing so is the only 'kind' or 'helpful' way to interact with this person. The rule watches for the specific phrasing tied to this scenario — the disease name and description of inverted emotional valence — combined with a request to flip normal response behavior.

Netzilo detection

Netzilo reports this behaviour when it is observed.

Signature ID
f1c8a53a-fc54-42b0-b442-ed5b2b4e356a
Severity
High

Why it matters

If successful, this bypasses the AI's content refusals and gets it to output material it's designed to withhold, using a sympathy-based social engineering pretext rather than a technical exploit.

What you can do

  • Treat claims of a disability or medical condition as a reason to be more careful, not a reason to disable safety behavior — verify through legitimate accessibility channels instead of trusting inline prompt text.
  • Review AI system prompts and guardrails to ensure no user-supplied condition or context can instruct the model to invert its refusal logic.
  • Log and review conversations where users assert unusual medical or cognitive conditions immediately before requesting normally-restricted content.
  • Distinguish this from legitimate medical, research, or fictional discussion of emotional-valence perception — the concerning pattern is the combination of the disease claim with a demand to bypass safe responses.

Known benign look-alikes

  • Medical literature: "Patients with inversed emotional valence perception may misinterpret neutral stimuli" (legitimate psychology/neuroscience context without jailbreak request)
  • NLP research: "The model reads text in its inversed sentiment polarity to serve as a baseline classifier" (technical ML context without disease framing or harmful-content mandate)
  • Creative writing: "The character suffers from a rare affliction that inverts her emotional reactions to stimuli" (narrative character description without AI safety bypass)

References

Related threats