Severity by source
CVSS:4.0/AV:N/AC:L/AT:N/PR:L/UI:N/VC:N/VI:N/VA:L/SC:N/SI:N/SA:N/E:X/CR:X/IR:X/AR:X/MAV:X/MAC:X/MAT:X/MPR:X/MUI:X/MVC:X/MVI:X/MVA:X/MSC:X/MSI:X/MSA:X/S:X/AU:X/R:X/V:X/RE:X/U:X
Network-reachable authenticated API with low complexity; availability set to High because the advisory describes severe CPU exhaustion causing denial of service, not mere degradation.
Primary rating from Vendor (VulnCheck).
CVSS VectorVendor: VulnCheck
Lifecycle Timeline
3Blast Radius
ecosystem impact- 3 pypi packages depend on vllm (3 direct, 0 indirect)
Ecosystem-wide dependent count for version 0.6.3.
DescriptionCVE.org
vLLM versions >= 0.6.3 and < 0.9.0 contain multiple regular expression denial of service (ReDoS) vulnerabilities. Several regex patterns - in vllm/lora/utils.py, the phi4mini tool parser, and the OpenAI-compatible serving chat endpoint - are susceptible to catastrophic backtracking. An attacker submitting crafted input with nested or repeated structures can trigger severe CPU consumption and performance degradation, resulting in denial of service.
AnalysisAI
Regular expression denial of service in vLLM versions 0.6.3 through 0.8.x exposes three distinct attack surfaces - the LoRA utility module, the phi4mini tool parser, and the OpenAI-compatible chat endpoint - to catastrophic regex backtracking, causing severe CPU exhaustion and service-wide denial of service. Authenticated API consumers can submit crafted inputs with deeply nested or repeated structures (e.g., ((((a|)+)+)+)) to trigger unbounded processing in Python's backtracking NFA regex engine. No public exploit identified at time of analysis, though the GHSA advisory discloses the exact vulnerable patterns and example malicious inputs, substantially lowering the reproduction barrier for anyone with API access.
Technical ContextAI
vLLM is a high-throughput LLM inference engine commonly deployed as an OpenAI-compatible API server, covered by CPE cpe:2.3:a:vllm:vllm:*:*:*:*:*:*:*:*. All three vulnerabilities fall under CWE-1333 (Inefficient Regular Expression Complexity): r"\((.*?)\)\$?$" in vllm/lora/utils.py line 173 processes LoRA adapter names; r'functools\[(.*?)\]' with re.DOTALL in vllm/entrypoints/openai/tool_parsers/phi4mini_tool_parser.py line 52 processes model output; and r'.*"parameters":\s*(.*)' in vllm/entrypoints/openai/serving_chat.py line 351 processes chat message content. Python's standard re module uses a backtracking NFA that provides no catastrophic backtracking protection, meaning adversarially crafted inputs can induce exponential regex engine state-space exploration and unbounded CPU consumption on a single request.
RemediationAI
The primary remediation is to upgrade vLLM to version 0.9.0 or later, which resolves all three ReDoS patterns; the upstream fix is tracked in PR #18454 and commit 4fc1bf813ad80172c1db31264beaef7d93fe0601 (https://github.com/vllm-project/vllm/pull/18454). If an immediate upgrade is not feasible, the GHSA advisory recommends enforcing explicit input length limits before any affected regex is evaluated - apply a hard maximum length to LoRA adapter name strings, to model_output content reaching the phi4mini parser, and to chat message content processed by serving_chat.py; this trades off some flexibility in accepted inputs but eliminates the exponential backtracking path. Disabling LoRA adapter support entirely (if unused in the deployment) eliminates the lora/utils.py attack surface with no inference capability impact. Restricting API access to trusted, verified consumers using network-layer controls (firewall rules, authentication middleware) reduces exposure consistent with the PR:L exploitation prerequisite. Full vendor advisory: https://github.com/vllm-project/vllm/security/advisories/GHSA-j828-28rj-hfhp.
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Rated critical severity (CVSS 10.0
vllm-project vllm version v0.6.2 contains a vulnerability in the MessageQueue.dequeue() API function. Rated critical sev
Information exposure in vLLM inference engine versions 0.8.3 to before 0.14.1. Invalid image requests to the multimodal
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Rated high severity (CVSS 7.5), th
vLLM before version 0.14.1 contains a server-side request forgery vulnerability in the MediaConnector class where incons
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Rated medium severity (CVSS 6.5),
Uncontrolled resource consumption in vLLM's OpenAI-compatible completions endpoint allows any authenticated API client t
Vllm versions up to 0.12.0 is affected by allocation of resources without limits or throttling (CVSS 6.5).
Race condition in vLLM's prompt embedding loader allows concurrent API requests to bypass the sparse tensor invariant gu
Remote code execution in vLLM 0.10.1 through 0.13.x lets an attacker who controls the model repository or path run arbit
Server-Side Request Forgery in vLLM's multimodal MediaConnector allows remote attackers to coerce the inference server i
Denial of service in vllm 0.19.0's OpenAI-compatible serving path allows remote unauthenticated attackers to exhaust sch
Same technique Denial Of Service
View allVendor StatusVendor
Share
External POC / Exploit Code
Leaving vuln.today
EUVD-2025-210290