# vLLM Malicious Model Weights Deserialization RCE

- **Severity:** High
- **Signature ID:** `7c1e9a34-5f2b-4d68-9a01-3e6b8f2c4d17`

## Summary

An attacker embeds a hidden pickle payload inside a model checkpoint file hosted on a model hub. When a vLLM inference server loads that checkpoint, the payload runs automatically inside the server process, giving the attacker code execution on the host.

## How the attack works

A malicious model file is prepared with a pickle __reduce__ trick, disguised as normal model weights. A victim's vLLM/PyTorch pipeline pulls the checkpoint and calls torch.load (with weights_only=False) or the vLLM weight iterator to deserialize it. Unpickling silently executes the embedded payload as part of loading, inside the same process that runs inference. The payload typically launches a shell, a file downloader, or an inline Python/interpreter one-liner to gain further access. The detection confirms this by tracing the exact process spawned from the model-loading step, following its real parent-child chain, and checking that the surrounding activity (checkpoint read, remote fetch, outbound connection) happened within the same short window and same process lineage — not just anywhere on the machine.

## Why it matters

An attacker can gain arbitrary code execution inside an AI inference server simply by getting it to load a poisoned model checkpoint — no direct network exploit needed. From there they can steal data, pivot into the host or cluster, or use the server as a foothold, since inference servers often run with access to GPUs, internal networks, or cloud credentials.

## What you can do

- Only load model weights from sources you trust and control; verify checksums or signatures before deployment.
- Load checkpoints with torch.load(weights_only=True) or use safetensors format instead of pickle-based weight files wherever possible.
- Run inference servers with least-privilege service accounts and restrict outbound network access from the process, so a payload can't easily download further tools.
- Review startup scripts and entrypoints that fetch weights at runtime, and avoid piping downloaded content directly into a shell or writing it to writable temp directories.

## Known benign look-alikes

- Container or Kubernetes entrypoints that wrap vLLM startup in "sh -c" and use curl/wget to stage weights from an internal registry before serving. These only reach the payload signature if the download is piped into a shell or written to /tmp, /var/tmp or /dev/shm and chmod +x'd.
- Distributed launchers (torchrun, Ray worker bootstrap, DeepSpeed) that spawn "python -c" one-liners importing socket or subprocess for rendezvous. This is the highest-noise benign case and is why the rule only reports.
- Interactive debugging in a Jupyter/IPython kernel that is itself a descendant of the vLLM process, where an engineer runs base64 or curl from the same lineage.
- CUDA kernel JIT compilation via torch.utils.cpp_extension (ninja, nvcc, gcc). These spawn shells but never match the payload signatures.

## References

- https://cwe.mitre.org/data/definitions/502.html
- https://pytorch.org/docs/stable/generated/torch.load.html
- https://attack.mitre.org/techniques/T1195/001/

---
Source: https://www.netzilo.com/threats/vllm-pickle-weights-rce
