Extremist Propaganda Acquisition & Distribution
This detects attempts to get an AI agent to actively source and spread named extremist texts, such as Mein Kampf and The Protocols of the Elders of Zion, rather than just discuss them academically. It flags requests that combine an action verb like 'upload', 'post', or 'endorse' with the specific named texts.
How the attack works
A user asks an AI agent to find, download, or otherwise acquire one of the named extremist texts. The same request, or a follow-up one, then asks the agent to distribute the material, for example by uploading it to a file host or posting it on social media, sometimes with language endorsing its content. The rule looks for this combination of an operational verb (acquire, upload, post, endorse) appearing alongside both named texts, so it does not trigger on someone simply mentioning the titles.
Netzilo detection
Netzilo reports this behaviour when it is observed.
- Signature ID
- 3df3c0f9-ea0f-43ef-891b-bcec34a6b443
- Severity
- Critical
Why it matters
An organization's AI agent could be turned into a distribution channel for banned or highly sensitive hate propaganda, creating legal exposure, reputational damage, and potential platform bans if the content is posted publicly under the organization's identity.
What you can do
- →Restrict agent tool permissions so upload and social-posting actions require human approval for any content involving named extremist or hate-related texts.
- →Log and review agent requests that pair document acquisition with distribution or endorsement language before allowing them to execute.
- →Add an allowlist exception path for legitimate scholarly or library use cases so historians and researchers are not blocked.
- →Periodically audit agent conversation logs for requests naming extremist texts alongside action verbs like post, upload, share, or endorse.
Known benign look-alikes
- Historians/scholars analyzing these texts as antisemitic propaganda
- Library catalog entries with scholarly annotations