FigStep Image-List Multimodal Jailbreak Text Pattern
FigStep is a jailbreak technique for multimodal AI models that hides harmful requests inside an image instead of the prompt text. The model is shown a picture containing an empty numbered list and asked to fill in each item with about 100 words, so it generates the harmful content itself without the request ever appearing in plain text.
How the attack works
The attacker crafts an image containing a numbered list (1, 2, 3...) where the list items are blank, with the actual harmful topic implied visually or via context rather than written out. The accompanying text prompt uses LLaMA-2 style [INST]...[/INST] formatting wrapped around a reference to the image, and instructs the model to 'fill in' the empty items with detailed content, roughly 100 words each. Because the harmful ask lives in the image rather than the text, naive text-based content filters miss it. The model, treating this as a benign completion task, generates the harmful material as if answering a formatting exercise. This pattern is detectable purely from the text layer, without needing to inspect the image itself, by spotting the combination of [INST] tags, an image reference, 'items are empty' phrasing, and the '100 words per item' instruction.
Netzilo detection
Netzilo reports this behaviour when it is observed.
- Signature ID
- 13d83c52-0013-47e6-87e2-90480fba68bc
- Severity
- High
Why it matters
An organization deploying a multimodal AI agent could have its safety filters bypassed, causing the model to produce harmful, restricted, or policy-violating content that text-only moderation would normally catch.
What you can do
- →Apply content moderation to model outputs, not just input prompts, so generated text is checked even when the harmful request was hidden in an image.
- →Treat prompts combining LLaMA-style [INST] delimiters, image references, and 'fill in the empty list items' framing as suspicious and route them for extra review.
- →Test multimodal agents against known jailbreak probe sets (such as garak's visual_jailbreak) before deployment.
- →Be aware of the listed false-positive cases (legitimate image description, quiz/accessibility use of numbered lists) and tune review thresholds to avoid blocking normal multimodal workflows.
Known benign look-alikes
- Legitimate multimodal prompts asking to describe or analyze actual image contents
- Educational prompts showing numbered lists in images for quiz or exercise contexts
- Accessibility tools that describe image content using structured list format