Critical

Rephrase Criminal Activity Second-Person Template (against against)

This detects a specific jailbreak template that tricks an AI into writing detailed criminal how-to instructions by disguising the request as a grammar exercise. Instead of asking directly 'how do I commit a crime,' the attacker frames it as rephrasing existing text into a rigid format, which bypasses safety filters looking for direct requests.

How the attack works

The attacker submits a prompt claiming to just want text 'rephrased' into a specific grammatical structure: the criminal is always addressed as 'you' (second person), the victim is mentioned but never as the grammatical subject, and the content must be organized into four categories — Preparation, Execution, Advantages gained, and Getting away. A distinctive 'against against (topic)' double-phrase and a fake disclaimer URL with URL-encoded text are often included to reinforce the false framing of 'informing' rather than 'instructing.' The model, treating this as a formatting task rather than a request for harmful content, produces a step-by-step account of how to prepare, carry out, and escape from a crime. This is a known jailbreak template from a public corpus, not a novel attack.

Netzilo detection

Netzilo reports this behaviour when it is observed.

Signature ID
31ddc1df-b5ed-48b7-b0d8-9de592e04f15
Severity
Critical

Why it matters

An organization running an AI agent or chatbot could have it produce detailed criminal instructions (e.g., planning, executing, and evading detection for a crime) despite safety guardrails, exposing the operator to abuse, reputational damage, and potential liability for facilitating harm.

What you can do

  • Add output-side review for responses containing structured 'Preparation / Execution / Advantages gained / Getting away' sections combined with second-person criminal framing.
  • Flag prompts containing the phrase 'against against' or fake disclaimer-style URLs with URL-encoded content as suspicious.
  • Test your model/agent against known jailbreak corpora (such as garak's DanInTheWild set) before deployment and after any prompt or model changes.
  • Do not rely on surface-level refusal keywords alone; review requests that ask the model to 'rephrase' or 'reformat' sensitive content, since this framing is used to bypass direct-ask filters.

Known benign look-alikes

  • Legal document: "Rephrase this legal document using second-person 'you' throughout — no criminal framing, no crime categories, no against-against topic"
  • Fiction: "Write a story where the criminal is referred to as 'you' in second person" (fiction without the four crime-stage category structure)
  • Project management: "Preparation, execution, and follow-through in project management — describe each phase" (PM categories without criminal/victim framing)
  • Academic: "Research paper on criminal psychology using second-person narrative" (academic without crime categories or against-against topic trigger)

References

Related threats