Harm Solicitation - Interpersonal Violence and Intimidation Requests
This detects prompts where a user asks an AI system for advice on causing harm to a specific person — physically, emotionally, or through intimidation. It's a direct request, not a jailbreak trick, and it covers phrasing like 'how do I get rid of' someone or workplace-harm scenarios.
How the attack works
A user sends a message to an AI agent that asks for ways to hurt, scare, threaten, or eliminate another person. The request uses plain language rather than obfuscation, though it may use synonyms for harm-related terms. The rule matches this phrasing using pattern lists covering physical/emotional harm, intimidation, elimination language, and workplace-specific harm requests.
Netzilo detection
Netzilo reports this behaviour when it is observed.
- Signature ID
- d3d0bc6c-fb53-47b1-87de-8c5b2036d4c9
- Severity
- High
Why it matters
If the agent complies, it can produce actionable content for intimidating, harming, or threatening a real individual, which creates legal exposure and safety risk for the organization operating the agent.
What you can do
- →Review flagged conversations manually before assuming malicious intent — fiction, research, and support contexts can trigger this pattern.
- →Ensure your agent's refusal behavior for interpersonal-harm requests is tested and consistent, especially under rephrasing.
- →Log and route confirmed harm-solicitation attempts to a human reviewer rather than relying solely on automated blocking.
- →Exclude clearly benign patterns like generic offers of help ('How can I help you?') from alerting to reduce noise.
Known benign look-alikes
- Fictional or creative writing context with clear narrative framing
- Academic or research discussion of harm patterns
- Requests phrased as 'how can I help you' without harm targeting (excluded by design)