Medium

Fictional Generation Demand Containing Target Output

Attackers ask an AI model to 'write a movie title' or 'compose a song lyric' but secretly demand the output contain an exact target string. Wrapping the real payload in a creative-writing request makes it look like harmless content rather than an instruction, which can slip past filters designed to catch direct commands.

How the attack works

The attacker opens with a benign-looking creative prompt, such as asking for a movie title, book name, or song chorus. Buried in that request is a constraint forcing the output to contain, append, or exactly match a specific target string, for example a leaked phrase, a phishing message, or a confirmation word. Because the request reads as fiction or entertainment, guardrails that look for direct commands often let it through. The model then produces the attacker's chosen string embedded in a 'creative' output, which can be used to prove a jailbreak worked, exfiltrate a marker, or plant text that a downstream system will trust.

Netzilo detection

Netzilo reports this behaviour when it is observed.

Signature ID
c65d2da8-1084-4c5c-8471-63f66e8f4846
Severity
Medium

Why it matters

An organization's AI safeguards can be bypassed to force the model to output attacker-chosen text, which may be used to confirm successful prompt injection, leak internal markers, or inject malicious strings into content later consumed by other automated systems.

What you can do

  • Treat prompts combining creative-generation requests with exact-match, quote, or append constraints as suspicious, not just prompts with direct commands.
  • Review outputs from creative-writing style prompts for suspicious embedded strings, links, or instructions before passing them to downstream systems.
  • Add output-side checks that flag exact string matches to known sensitive or attacker-controlled values, not just input-side keyword filtering.
  • Log and monitor for repeated creative-prompt patterns pivoting from a benign prefix into a strict output constraint, as this shape indicates probing for jailbreak success.

Known benign look-alikes

  • Genuine creative-writing requests without exact-output constraint.
  • Educational prompts about generative creative tasks.
  • Translation of existing creative content for legitimate review.

References

Related threats