High

Malicious Skill Update or Mutation

An add-on tool (skill) for an AI agent can be reviewed and approved while it's harmless, then quietly updated later to steal data, plant a backdoor, or manipulate the AI through prompt injection. Because trust was granted before the malicious change, the update can slip past initial security checks.

How the attack works

An attacker (or a compromised publisher account) first publishes a benign, useful skill and gets it approved and installed by an AI agent platform. Once the skill is trusted and in use, the attacker pushes a version update or re-registers the skill with new code. This update introduces harmful behavior such as exfiltrating data through tool responses, embedding a backdoor, or injecting hidden instructions that manipulate the agent. The malicious change is detected by watching for suspicious patterns in what the skill returns or requests right after a version change or re-registration event.

Netzilo detection

Netzilo reports this behaviour when it is observed.

Signature ID
a59db02b-22b3-48f5-9219-0dad1df6f1a8
Severity
High

Why it matters

An organization can lose data, have its AI agent's behavior silently hijacked, or have a backdoor persist inside a trusted workflow, all while believing the skill was already vetted and safe.

What you can do

  • Require manual re-review and re-approval whenever a skill's version changes, not just at first install.
  • Pin skills to specific reviewed versions and alert on any auto-update or re-registration.
  • Monitor tool responses and arguments for unexpected data patterns, especially right after an update.
  • Track skill publisher identity and account changes to catch compromised or hijacked publisher accounts.

Known benign look-alikes

  • Legitimate skill updates that require re-authorization
  • Platform migration notices from verified publishers
  • Dependency update notifications in development environments

References

Related threats