Severity by source
AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H
Network-accessible API requires authenticated low-privilege credential; only availability is impacted via unchecked resource fan-out; no scope change or confidentiality/integrity loss.
Primary rating from Vendor (GitHub_M).
CVSS VectorVendor: GitHub_M
Lifecycle Timeline
3Blast Radius
ecosystem impact- 1 pypi packages depend on vllm (1 direct, 0 indirect)
Ecosystem-wide dependent count for version 0.19.0.
DescriptionCVE.org
vLLM is an inference and serving engine for large language models. From 0.19.0 until 0.26.0, the /v1/completions CompletionRequest.prompt field in vllm/entrypoints/openai/completion/protocol.py accepts an unbounded list[str] or list[list[int]], prompt_to_seq() in vllm/renderers/inputs/preprocess.py and OnlineRenderer.preprocess_completion() in vllm/renderers/online_renderer.py expand every element, and vllm/entrypoints/openai/completion/serving.py creates one engine generator and response slot per prompt, allowing an authenticated API client to exhaust CPU, memory, async scheduling capacity, engine request slots, and response buffering with one request. This issue is fixed in version 0.26.0.
AnalysisAI
Uncontrolled resource consumption in vLLM's OpenAI-compatible completions endpoint allows any authenticated API client to exhaust CPU, memory, async scheduling capacity, and engine request slots with a single crafted request. Affected versions span 0.19.0 through 0.25.x; the issue is fixed in 0.26.0. No public exploit code has been identified and CISA has not listed this in KEV, but the attack is trivially constructable by any client with a valid API credential.
Technical ContextAI
vLLM (cpe:2.3:a:vllm-project:vllm) is a high-throughput inference and serving engine for large language models that exposes an OpenAI-compatible HTTP API. The root cause (CWE-400, Uncontrolled Resource Consumption) lies in three cooperating components: CompletionRequest.prompt in vllm/entrypoints/openai/completion/protocol.py accepts a list[str] or list[list[int]] with no upper-bound validator; prompt_to_seq() in vllm/renderers/inputs/preprocess.py and OnlineRenderer.preprocess_completion() in vllm/renderers/online_renderer.py unconditionally expand every list element into a separate inference sequence; and vllm/entrypoints/openai/completion/serving.py allocates one engine generator and one response buffer slot per expanded prompt. Because the engine's async scheduler and response-buffering are both driven by this per-prompt fan-out with no rate or size check at the protocol layer, a single API call with an arbitrarily long prompt list can saturate all four resource dimensions simultaneously.
RemediationAI
Upgrade to vLLM 0.26.0, which introduces a bounded validator for the prompt field in CompletionRequest via the VLLM_MAX_COMPLETION_PROMPTS environment variable (see PR #47845 and commit 675f4295cdfe0d870471c2b51bfeca3a68a9569e). The fix rejects requests whose prompt list length exceeds the configured maximum before any engine allocation occurs. If an immediate upgrade is not feasible, operators can restrict the /v1/completions endpoint to highly trusted clients via network policy or API gateway rate-limiting on request body size, and should consider setting per-client concurrency limits at the reverse-proxy layer; however, these controls do not fix the underlying unbounded fan-out and carry the trade-off of reduced API accessibility. The vendor advisory is at https://github.com/vllm-project/vllm/security/advisories/GHSA-87x5-vmc3-756j.
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Rated critical severity (CVSS 10.0
vllm-project vllm version v0.6.2 contains a vulnerability in the MessageQueue.dequeue() API function. Rated critical sev
Information exposure in vLLM inference engine versions 0.8.3 to before 0.14.1. Invalid image requests to the multimodal
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Rated high severity (CVSS 7.5), th
vLLM before version 0.14.1 contains a server-side request forgery vulnerability in the MediaConnector class where incons
Out-of-memory worker crashes in vLLM can be induced by a single small compressed audio payload submitted to the /v1/chat
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Rated medium severity (CVSS 6.5),
GPU memory exhaustion in vLLM versions prior to 0.28.0 lets an attacker with low-privileged API access (PR:L in the CVSS
Multimodal inference requests to vLLM deployments running versions prior to 0.28.0 can trigger a process-level assertion
Vllm versions up to 0.12.0 is affected by allocation of resources without limits or throttling (CVSS 6.5).
Race condition in vLLM's prompt embedding loader allows concurrent API requests to bypass the sparse tensor invariant gu
Remote code execution in vLLM 0.10.1 through 0.13.x lets an attacker who controls the model repository or path run arbit
Same weakness CWE-400 – Uncontrolled Resource Consumption
View allSame technique Denial Of Service
View allVendor StatusVendor
SUSE
Severity: ModerateShare
External POC / Exploit Code
Leaving vuln.today
EUVD-2026-58067
GHSA-87x5-vmc3-756j