High

NSFW Character Sheet Generation Unlock

This detects a jailbreak trick where someone hides a fake permission notice inside a fictional character creation template to trick an AI into generating explicit content. Instead of directly asking for NSFW material, the attacker frames it as already-approved by embedding a phrase like "(NSFW generations are allowed)" alongside requests for fetish or kink details as part of the character's persona.

How the attack works

The attacker submits a structured roleplay character sheet template — the kind used for collaborative fiction or game-style AI personas. Inside the template, they insert a parenthetical or inline statement declaring that NSFW generation is already allowed or enabled, phrased as a fact rather than a request. The same template also asks for sexual or explicit personality traits such as fetishes or kinks as normal character fields. By combining a false permission grant with routine-looking persona fields, the attacker tries to make the AI treat explicit content generation as already authorized, bypassing its content moderation.

Netzilo detection

Netzilo reports this behaviour when it is observed.

Signature ID
b80c0c4b-9438-4a62-a8a1-eaf12d01ff99
Severity
High

Why it matters

An AI agent can be tricked into producing sexual or explicit content it would otherwise refuse, undermining content moderation controls and exposing the operator to policy, legal, or reputational risk from generated material.

What you can do

  • Review character-creation or roleplay templates before they reach the model for embedded permission statements about NSFW or restricted content.
  • Do not let inline text in user-submitted templates be treated as system-level policy; keep content permissions controlled outside user input.
  • Flag prompts that combine explicit persona attributes (fetishes, kinks) with statements claiming content restrictions are lifted or already approved.
  • Distinguish moderation-language false positives (e.g. "NSFW filtering is applied") from actual unlock phrases (e.g. "NSFW generations are allowed") when tuning detection.

Known benign look-alikes

  • "NSFW content filtering is applied to all characters" — NSFW as a moderation/blocking term, not a permission grant
  • "This platform uses NSFW detection to screen character submissions" — detection/screening context, not generation unlock
  • "Character Sheet: Name, Class, Stats — no adult content allowed on this server" — character sheet without NSFW unlock

References

Related threats