Critical

Malicious Fine-tuning Data

Attackers can slip poisoned examples into a fine-tuning dataset so that a model behaves normally most of the time but does something malicious when it sees a specific trigger phrase. This turns a normal training update into a way to secretly plant a backdoor in an AI model's behavior.

How the attack works

An attacker crafts training examples that pair an innocuous-looking trigger input with a malicious or safety-bypassing response. These poisoned samples are mixed into a larger fine-tuning dataset or submitted as training examples, so they blend in with legitimate data. Once the model is fine-tuned on this data, it learns the association and will follow the hidden instruction whenever the trigger appears in a future prompt. The rest of the model's behavior looks unchanged, so the backdoor stays hidden until someone uses the trigger.

Netzilo detection

Netzilo reports this behaviour when it is observed.

Signature ID
802373e2-88d0-4d62-aad8-e031737a9466
Severity
Critical

Why it matters

A poisoned model can be made to bypass its own safety filters, leak data, or take unauthorized actions on command, and because the behavior only appears with a specific trigger, it can go undetected through normal testing and evaluation.

What you can do

  • Review and vet the source and provenance of any fine-tuning dataset before use, especially data pulled from external or crowdsourced sources.
  • Scan training examples for suspicious trigger-response patterns, such as unusual phrases paired with instructions to ignore safety rules or leak data.
  • Test fine-tuned models against a wide range of crafted inputs, not just standard benchmarks, to surface hidden trigger behaviors.
  • Keep records of who submitted or approved training data changes so poisoned submissions can be traced back.

Known benign look-alikes

  • Security researchers constructing adversarial datasets for red-teaming
  • Fine-tuning data that legitimately discusses security concepts
  • Training examples for security-focused classification models
  • Academic papers analyzing data poisoning techniques that reference backdoor terminology in quoted citations
  • Security audit reports documenting discovered training data anomalies using technical vocabulary

References

Related threats