Critical

Self-Harm and Eating Disorder Facilitation Request

This rule flags messages sent to an AI assistant that directly ask for methods to self-harm, induce vomiting, starve, maintain dangerously low-calorie diets, or hide these behaviors from others. It distinguishes these operational requests from legitimate mental health conversations.

How the attack works

A user sends a prompt to an AI agent asking a direct, actionable question rather than seeking support — for example asking what object to use for self-harm, how to make themselves throw up, how to develop anorexia, or how to hide scars or an eating disorder from family or friends. The rule looks for this specific pattern of operational, method-focused requests rather than general discussion of mental health struggles. If the assistant lacks proper safety guardrails, it may respond with the requested harmful information instead of refusing or redirecting to support resources.

Netzilo detection

Netzilo reports this behaviour when it is observed.

Signature ID
035c4a17-a3c1-4735-a87c-d5d7631b55c5
Severity
Critical

Why it matters

An organization deploying an unguarded AI assistant risks it directly facilitating self-harm or disordered eating in a vulnerable user, creating serious safety and liability exposure and potential reputational and legal harm.

What you can do

  • Configure the AI assistant to refuse operational self-harm and eating disorder requests and redirect to crisis resources instead of answering.
  • Review logs for this pattern to identify at-risk users and route them to human support rather than treating this purely as a security event.
  • Exclude legitimate clinical, training, or academic content from alerting by checking context before escalating.
  • Test the assistant's refusal behavior against these prompt patterns before deployment and after any model or prompt updates.

Known benign look-alikes

  • Healthcare providers documenting patient-reported behaviors in clinical notes
  • Mental health crisis hotline training materials listing at-risk phrases
  • Academic research on self-harm prevention quoting harmful prompt examples
  • Legitimate medical questions about eating disorder treatment and recovery

References

Related threats