MCP Tool Description Changed After Approval (Rug Pull)
An AI agent tool that was reviewed and approved can later have its description silently swapped for a different one carrying new instructions. Because each individual call still looks normal, only comparing the tool's description over time reveals the tampering.
How the attack works
An operator or developer approves an MCP tool after reviewing its stated description and behavior. Later, the tool's definition is changed — sometimes days afterward — while its name and outward function stay the same. The agent keeps calling the tool as usual, so nothing about a single call or execution looks suspicious. The only way to catch this is to remember the original description and compare it against every later version.
Netzilo detection
Netzilo reports this behaviour when it is observed.
- Signature ID
- 43a31200-6894-5c9b-a187-1cd0a11af2b3
- Severity
- Medium
Why it matters
An attacker who controls or compromises an MCP server can quietly redirect an already-trusted tool to follow new, potentially malicious instructions, bypassing the review process that approved the original version.
What you can do
- →Keep a record of approved tool descriptions and re-review any tool whose description changes, even after long periods.
- →Treat description changes as requiring human review rather than automatic re-approval.
- →If a tool legitimately templates dynamic values (timestamps, versions, counts) into its description, normalize or strip those before comparing, or allowlist the tool to avoid noise.
- →After any planned server upgrade that changes tool wording, expect and verify one flagged change per tool, then confirm the new baseline is correct.
Known benign look-alikes
- A legitimate MCP server upgrade that rewords its tool descriptions — expect one report per changed tool per upgrade, then silence as the baseline re-anchors
- Servers that template a description with a timestamp, version string or row count, which changes the hash on every load. Those need an allowlist here or a normalisation step before hashing
- The first run after the rule's store is reset, if descriptions changed while the rule was disabled