Severity by source
AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H
Network-reachable inference API (AV:N), low-privilege request submitter to a shared service (PR:L), no user interaction, deterministic assertion crash gives A:H with no C/I impact.
Primary rating from Vendor (GitHub_M).
CVSS VectorVendor: GitHub_M
Lifecycle Timeline
3Blast Radius
ecosystem impact- 2 pypi packages depend on vllm (2 direct, 0 indirect)
Ecosystem-wide dependent count for version 0.28.0.
DescriptionCVE.org
vLLM is an inference and serving engine for large language models. Prior to 0.28.0, the default mirrored multimodal LRU cache can commit a media hash in the frontend sender cache during multimodal rendering and before engine admission, while the engine receiver cache never receives the payload if that request is rejected. A later request reusing the same media hash causes MultiModalProcessorSenderCache to send no payload and MultiModalReceiverCache to reach an assertion with the message "Expected a cached item," producing a shared-service availability failure. This issue is fixed in version 0.28.0.
AnalysisAI
Multimodal inference requests to vLLM deployments running versions prior to 0.28.0 can trigger a process-level assertion failure that takes down the shared serving engine, denying service to all concurrent users. The flaw is a state desynchronization in the default mirrored multimodal LRU cache: a media hash can be committed in the frontend MultiModalProcessorSenderCache during rendering before engine admission, but if the request is rejected, the engine-side MultiModalReceiverCache never stores the corresponding payload, so any later request reusing that hash reaches an assertion reading "Expected a cached item." Exploitation requires only the ability to submit requests to the service (PR:L, network-reachable, no user interaction), is confined to deployments serving multimodal models, and results in availability impact only; no public exploit code was identified at time of analysis and upgrading to 0.28.0 removes the issue.
Unlock full vulnerability intelligence
- Risk assessment & exploitation conditions
- Attack chain visualization
- Remediation with exact patch versions
- Threat intelligence from 22 sources
- Personal watchlist & email alerts
No credit card · 7-day full trial
Attack ChainAIDerived
Hypothetical attack flow derived from CVE metadata
Vulnerability AssessmentAI
| Exploitation | Requires a vLLM deployment (prior to 0.28.0) serving multimodal models with the default mirrored multimodal LRU cache in use. … Additional conditions and limiting factors are described in the full assessment. |
| Risk Assessment | This is a genuine but moderate-severity availability issue, not a critical priority. … Full risk analysis with EPSS, KEV, and SSVC signal comparison available after sign-in. |
| Exploit Scenario | Full exploit scenario with step-by-step reproduction available after sign-in. |
| Remediation | Upgrade to vLLM 0.28.0 or later, which is the vendor-released fix; the patched release tag is https://github.com/vllm-project/vllm/releases/tag/v0.28.0, the specific change is commit 396204230423b7cc6798300926b8fa30190d26a9, and the advisory is https://github.com/vllm-project/vllm/security/advisories/GHSA-ph3r-5jfg-f84f. … Detailed patch versions, workarounds, and compensating controls in full report. |
Threat intelligence, references, and detailed analysis are available after sign-in.
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Rated critical severity (CVSS 10.0
vllm-project vllm version v0.6.2 contains a vulnerability in the MessageQueue.dequeue() API function. Rated critical sev
Information exposure in vLLM inference engine versions 0.8.3 to before 0.14.1. Invalid image requests to the multimodal
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Rated high severity (CVSS 7.5), th
vLLM before version 0.14.1 contains a server-side request forgery vulnerability in the MediaConnector class where incons
Out-of-memory worker crashes in vLLM can be induced by a single small compressed audio payload submitted to the /v1/chat
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Rated medium severity (CVSS 6.5),
GPU memory exhaustion in vLLM versions prior to 0.28.0 lets an attacker with low-privileged API access (PR:L in the CVSS
Uncontrolled resource consumption in vLLM's OpenAI-compatible completions endpoint allows any authenticated API client t
Vllm versions up to 0.12.0 is affected by allocation of resources without limits or throttling (CVSS 6.5).
Race condition in vLLM's prompt embedding loader allows concurrent API requests to bypass the sparse tensor invariant gu
Remote code execution in vLLM 0.10.1 through 0.13.x lets an attacker who controls the model repository or path run arbit
Same weakness CWE-617 – Reachable Assertion
View allSame technique Denial Of Service
View allShare
External POC / Exploit Code
Leaving vuln.today
EUVD-2026-92843
GHSA-ph3r-5jfg-f84f