High

vLLM Unbounded Fan-Out Parameter DoS Attempt

An AI agent (or something it runs) can send one specially crafted request to a vLLM inference server asking for an enormous number of completions in a single call. Because vLLM doesn't cap this number, the server tries to fulfill it, exhausts memory and CPU, and gets killed by the operating system — taking down the service for everyone else using it.

How the attack works

The attacker sends an OpenAI-compatible completion request to a vLLM endpoint, setting the 'n', 'best_of', or 'num_return_sequences' parameter to an extreme value instead of a normal small number. vLLM's request handler loops over this value, duplicating the request internally once per requested completion, with no upper limit enforced. This floods the server's single-threaded event loop, blocking all other connections, while memory usage grows by gigabytes per second. Eventually the operating system's out-of-memory killer terminates the server process, causing an outage for all users of that vLLM instance.

Netzilo detection

Netzilo reports this behaviour when it is observed.

Signature ID
7a2f4c18-9d63-4e5b-b1a7-0c8e5d34f962
Severity
High

Why it matters

A single malicious or misconfigured request can take down a shared vLLM inference server, denying service to every other agent or user relying on it, with no advance warning beyond the outbound request itself.

What you can do

  • Set and enforce a maximum allowed value for 'n', 'best_of', and 'num_return_sequences' in front of or inside your vLLM deployment, e.g. via a reverse proxy or request-validation layer.
  • Run vLLM instances with resource limits (memory cgroups, container limits) so an OOM event kills a contained process rather than the whole host.
  • Monitor for requests with unusually large fan-out parameters and distinguish one-off astronomical values (likely abuse) from steady, moderate-but-elevated values used by legitimate reranking or best-of-N sampling pipelines.
  • If you knowingly run best-of-N or load-testing workloads against a vLLM instance, isolate them on dedicated hosts and document expected fan-out ranges so anomalies stand out.

Known benign look-alikes

  • Best-of-N / self-consistency sampling harnesses that legitimately request hundreds of completions from a dedicated local vLLM instance. Raise FANOUT_REPORT in the script for those hosts rather than allow-listing them.
  • Capacity or load testing of a self-hosted vLLM deployment (benchmark_serving.py, lm-eval-harness sweeps) driven by a platform engineer through an agent shell.
  • Reranking / candidate-generation pipelines that set best_of high on purpose; these show a steady cadence of identical requests rather than a single astronomical value, and are surfaced as "excessive" rather than "astronomical".
  • A JSON body that happens to contain an unrelated top-level key literally named "n" with a large numeric value; mitigated by requiring POST plus a completions-style URL path before the body is parsed at all.

References

Related threats