High

Secret Key Exfiltration Request

This rule flags messages where a user directly asks an AI agent to disclose a secret, remembered, or hidden key value. It looks for the key being named and then immediately followed by a request for its literal value, often disguised as a task-ending or role-reset instruction.

How the attack works

An attacker sends a message to the AI agent that names a secret, session, or system key. Right after naming it, they add a demand to see its exact value, sometimes phrasing it as 'give it verbatim' or 'without translation.' The message is often framed as an 'end of task' or role-reset instruction, trying to trick the agent into thinking the secret is now safe to share. This can also appear as a literal placeholder string like {$SECRET_KEY} that the attacker hopes the system will fill in and leak.

Netzilo detection

Netzilo reports this behaviour when it is observed.

Signature ID
ffe49cbb-b7e3-4c64-bca7-dc2aef3dc845
Severity
High

Why it matters

If successful, the agent discloses a system or session secret key to the attacker, which can be used to impersonate the system, access protected resources, or escalate further attacks.

What you can do

  • Review agent logs for prompts that combine key references with disclosure demands, especially those framed as task-ending or role-reset instructions.
  • Ensure agents are configured to never output secret or session key values regardless of how the request is phrased.
  • Separate secret values from any context the agent can read or repeat back to users.
  • Treat this as detection of one specific exfiltration attempt pattern, not full coverage of all prompt injection techniques.

Known benign look-alikes

  • Documentation referencing a bare SECRET_KEY environment variable name without an exfiltration cue
  • Configuration file snippets that set or read SECRET_KEY without requesting its value be disclosed
  • General discussion of secret-management best practices

References

Related threats