Severity by source
AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H
Description states unauthenticated exploitation by default (PR:N), remote and low-complexity, with pure availability impact (A:H) and no confidentiality/integrity effect; auth is a deployment-specific control.
Primary rating from Vendor (GitHub_M).
CVSS VectorVendor: GitHub_M
Lifecycle Timeline
3Blast Radius
ecosystem impact- 4 pypi packages depend on vllm (4 direct, 0 indirect)
Ecosystem-wide dependent count for version 0.24.0.
DescriptionCVE.org
vLLM is an inference and serving engine for large language models. Prior to 0.24.0, the input_audio handling path for /v1/chat/completions calls AudioMediaIO.load_bytes or AudioMediaIO.load_file without passing VLLM_MAX_AUDIO_DECODE_DURATION_S to the shared audio decoder. An unauthenticated client can therefore submit a small compressed audio input that expands into a very large float32 PCM allocation, bypassing the duration guard already used by /v1/audio/transcriptions and causing an out-of-memory worker crash. Inline data URLs reach this path without being bounded by VLLM_AUDIO_FETCH_TIMEOUT. The issue affects deployments serving an audio-capable model, and authentication changes only the deployment-specific reachability. This issue is fixed in version 0.24.0.
AnalysisAI
Out-of-memory worker crashes in vLLM can be induced by a single small compressed audio payload submitted to the /v1/chat/completions input_audio path, because that code path decodes audio without the VLLM_MAX_AUDIO_DECODE_DURATION_S guard already enforced on /v1/audio/transcriptions - an unbounded resource allocation flaw (CWE-770) that expands a tiny compressed file into a very large float32 PCM buffer. The issue affects every vLLM release prior to 0.24.0 that serves an audio-capable model and exposes the chat completions endpoint, with inline data URLs additionally escaping VLLM_AUDIO_FETCH_TIMEOUT; impact is availability-only (no confidentiality or integrity loss). …
Unlock full vulnerability intelligence
- Risk assessment & exploitation conditions
- Attack chain visualization
- Remediation with exact patch versions
- Threat intelligence from 22 sources
- Personal watchlist & email alerts
Free forever · No credit card required
Attack ChainAIDerived
Hypothetical attack flow derived from CVE metadata
Vulnerability AssessmentAI
| Exploitation | Requires a vLLM deployment (<0.24.0) that serves an audio-capable model and exposes the /v1/chat/completions endpoint accepting input_audio. … Additional conditions and limiting factors are described in the full assessment. |
| Risk Assessment | All signals point to a genuine but bounded availability-only issue. … Full risk analysis with EPSS, KEV, and SSVC signal comparison available after sign-in. |
| Exploit Scenario | Full exploit scenario with step-by-step reproduction available after sign-in. |
| Remediation | Vendor-released patch: upgrade to vLLM 0.24.0, which threads max_duration_s=envs.VLLM_MAX_AUDIO_DECODE_DURATION_S into AudioMediaIO.load_bytes and AudioMediaIO.load_file (PR 45908, commit 3d20275bb4d434f53055c3c0b645fd8bb072965e); the release is at https://github.com/vllm-project/vllm/releases/tag/v0.24.0 and the advisory at https://github.com/vllm-project/vllm/security/advisories/GHSA-hcwq-8wjf-3gcr. … Detailed patch versions, workarounds, and compensating controls in full report. |
Threat intelligence, references, and detailed analysis are available after sign-in.
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Rated critical severity (CVSS 10.0
vllm-project vllm version v0.6.2 contains a vulnerability in the MessageQueue.dequeue() API function. Rated critical sev
Information exposure in vLLM inference engine versions 0.8.3 to before 0.14.1. Invalid image requests to the multimodal
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Rated high severity (CVSS 7.5), th
vLLM before version 0.14.1 contains a server-side request forgery vulnerability in the MediaConnector class where incons
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Rated medium severity (CVSS 6.5),
Uncontrolled resource consumption in vLLM's OpenAI-compatible completions endpoint allows any authenticated API client t
Vllm versions up to 0.12.0 is affected by allocation of resources without limits or throttling (CVSS 6.5).
Race condition in vLLM's prompt embedding loader allows concurrent API requests to bypass the sparse tensor invariant gu
Remote code execution in vLLM 0.10.1 through 0.13.x lets an attacker who controls the model repository or path run arbit
Server-Side Request Forgery in vLLM's multimodal MediaConnector allows remote attackers to coerce the inference server i
Denial of service in vllm 0.19.0's OpenAI-compatible serving path allows remote unauthenticated attackers to exhaust sch
Same technique Denial Of Service
View allShare
External POC / Exploit Code
Leaving vuln.today
EUVD-2026-80880
GHSA-hcwq-8wjf-3gcr