vLLM is a focused inference-serving framework for large language models that has become prominent in machine-learning deployment pipelines, and its vulnerability footprint reflects the security challenges of exposing model-serving endpoints to untrusted input. Vulnerabilities affecting the vendor skew strongly toward critical-severity outcomes, concentrating across weakness classes that are characteristic of networked Python services handling complex user-supplied payloads: untrusted deserialization, uncontrolled resource allocation, input-validation gaps, server-side request forgery, and code-injection vectors. The narrow product scope means that a single vLLM deployment may aggregate multiple of these weakness classes, and the vendor's prominence in the ML-infrastructure ecosystem amplifies the downstream impact of flaws in model serving and API handling. Defenders deploying vLLM should treat framework updates as security-critical, isolate model-serving endpoints, and apply strict input validation and resource throttling at the application layer; live severity, exploitation, and exposure counts are shown alongside this summary.
The number and severity of CVEs published that impact products developed by Vllm over time
Signals from CVEs in this vendor scope (51 CVEs).
51 CVEs · Highest risk first
| CVE | Published | CVSS | Risk | KEV | Exploit |
|---|---|---|---|---|---|
CVE-2026-22778CRITICAL vLLM is an inference and serving engine for large language models (LLMs). From 0.8.3 to before 0.14.1, when an invalid image is sent to vLLM's multimodal endpoint, PIL throws an er | Feb 2, 2026 | 9.8 | 53 | NO | YES |
CVE-2026-48746CRITICAL vLLM is an inference and serving engine for large language models (LLMs). From 0.3.0 until 0.22.0, a vulnerability in ASGI web servers and starlette's trust on those web servers en | Jun 22, 2026 | 9.1 | 39 | NO | NO |
CVE-2026-54236MEDIUM vLLM is an inference and serving engine for large language models (LLMs). Prior to 0.23.1rc0, the fix for CVE-2026-22778, which introduced a sanitize_message helper that strips obj | Jun 22, 2026 | 5.3 | 38 | NO | YES |
CVE-2026-56340HIGH vLLM versions >= 0.10.2 and < 0.13.0 are missing sparse tensor validation in multimodal embeddings processing. Because PyTorch disables sparse tensor invariant checks by default, a | Jun 20, 2026 | 8.8 | 37 | NO | NO |
CVE-2026-22807CRITICAL vLLM is an inference and serving engine for large language models (LLMs). Starting in version 0.10.1 and prior to version 0.14.0, vLLM loads Hugging Face `auto_map` dynamic modules | Jan 21, 2026 | 9.8 | 37 | NO | NO |
CVE-2026-54232HIGH vLLM is an inference and serving engine for large language models (LLMs). Prior to 0.22.1, the vLLM Dockerfile is vulnerable to a dependency confusion attack through the flashinfer | Jun 22, 2026 | 8.8 | 36 | NO | NO |
CVE-2026-41523HIGH vLLM is an inference and serving engine for large language models (LLMs). Prior to 0.22.0, an assert-based security check in vLLM's activation function loading allows any unauthent | Jun 22, 2026 | 7.5 | 34 | NO | NO |
CVE-2026-5497HIGH vLLM versions 0.8.0 and later are vulnerable to an Out-of-Memory (OOM) Denial of Service (DoS) attack due to unbounded frame count processing in the `VideoMediaIO.load_base64()` me | Jun 11, 2026 | 7.5 | 34 | NO | NO |
CVE-2026-55574HIGH vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Prior to 0.24.0, the structured_outputs.regex API parameter passes a user-supplied regular exp | Jul 6, 2026 | 7.5 | 33 | NO | NO |
CVE-2026-54234HIGH vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Prior to 0.24.0, a frontend-legal multi-request speculative decoding workload can cause the re | Jul 6, 2026 | 7.5 | 33 | NO | NO |
Signals from CVEs in this vendor scope (51 CVEs).
An overview of all social media posts that mention a CVE ID that affects a product developed by Vllm.
Media articles that mention a CVE ID that affects a product developed by Vllm — matched by CVE ID, not by vendor name.