Skip to main content

NVIDIA TensorRT-LLM EUVDEUVD-2026-44470

| CVE-2026-24271 MEDIUM
Allocation of Resources Without Limits or Throttling (CWE-770)
2026-07-14 nvidia GHSA-3wc9-fv9q-6v3m
6.2
CVSS 3.1 · Vendor: nvidia
Share

Severity by source

Vendor (nvidia) PRIMARY
6.2 MEDIUM
AV:L/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H
vuln.today AI
7.5 HIGH

OpenAI-compatible inference API is network-facing by design; AV:N and PR:N reflect unauthenticated network access, with availability-only impact.

3.1 AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H
4.0 AV:N/AC:L/AT:N/PR:N/UI:N/VC:N/VI:N/VA:H/SC:N/SI:N/SA:N

Primary rating from Vendor (nvidia).

CVSS VectorVendor: nvidia

Attack Vector
Local
Attack Complexity
Low
Privileges Required
None
User Interaction
None
Scope
Unchanged
Confidentiality
None
Integrity
None
Availability
High

Lifecycle Timeline

1
Analysis Generated
Jul 22, 2026 - 13:32 vuln.today

DescriptionCVE.org

NVIDIA TensorRT-LLM contains a vulnerability in the OpenAI-compatible inference API, where an attacker could cause allocation of GPU resources without limits or throttling. A successful exploit of this vulnerability might lead to denial of service.

AnalysisAI

Unbounded GPU resource allocation in NVIDIA TensorRT-LLM's OpenAI-compatible inference API allows an attacker to exhaust GPU resources on the host system, resulting in denial of service. Affected versions span all releases up to and including v1.3.0 rc14. No public exploit or active exploitation has been identified at time of analysis; SSVC classifies exploitation as none and technical impact as partial, consistent with a localized availability-only impact.

Technical ContextAI

NVIDIA TensorRT-LLM (cpe:2.3:a:nvidia:tensorrt-llm:*:*:*:*:*:*:*:*) is a high-performance LLM inference optimization library that exposes an OpenAI-compatible HTTP API for serving large language models on NVIDIA GPU hardware. The root cause is CWE-770 (Allocation of Resources Without Limits or Throttling): the inference API does not enforce quotas, caps, or rate limits on GPU memory or compute resource allocation per request or per client. A caller can submit requests that drive unbounded GPU allocations, starving other workloads or crashing the inference server. The official CVSS vector records AV:L, suggesting the API may be treated as a locally-scoped service, though OpenAI-compatible APIs are architecturally network-facing; see confidence notes.

RemediationAI

Upgrade NVIDIA TensorRT-LLM to a version beyond v1.3.0 rc14; the exact patched release version was not independently confirmed from the available references - consult the NVIDIA security bulletin and TensorRT-LLM release notes directly for a confirmed fix version. If upgrading is not immediately possible, apply compensating controls at the API layer: configure an API gateway or reverse proxy (e.g., nginx) in front of the inference endpoint with per-client and global request rate limits and concurrency caps to restrict unbounded resource allocation. Additionally, restrict network access to the inference API to trusted internal IP ranges only, reducing the attacker population. Note that rate limiting at the proxy layer does not eliminate the root CWE-770 flaw but significantly raises the effort required for exploitation. Monitor GPU memory utilization and set alerting thresholds to detect anomalous allocation patterns early.

CVE-2026-24142 CRITICAL
9.8 May 20

Deserialization of untrusted data in NVIDIA TensorRT-LLM across all platforms allows a local, low-privileged attacker to

CVE-2026-24163 CRITICAL
9.8 May 20

Unsafe deserialization in NVIDIA TensorRT-LLM's RPC testing component allows a local high-privileged attacker to trigger

CVE-2025-33255 CRITICAL
9.8 May 20

Unsafe deserialization in NVIDIA TensorRT-LLM's MPI server component allows a high-privileged local attacker to achieve

CVE-2026-24233 HIGH
8.4 Jul 14

Insecure deserialization in NVIDIA TensorRT-LLM for Linux lets a local, low-privileged attacker abuse a weakness in the

CVE-2026-47472 HIGH
7.8 Jul 14

Local privilege-context deserialization in NVIDIA TensorRT-LLM lets an attacker who already has same-user access to a ho

CVE-2026-47471 HIGH
7.5 Jul 14

Heap-based buffer overflow in NVIDIA TensorRT-LLM's tensor deserialization path lets an adjacent, unauthenticated attack

CVE-2026-24160 HIGH
7.5 May 20

Null pointer dereference in NVIDIA TensorRT-LLM across all supported platforms allows a local attacker to crash the appl

CVE-2026-47473 HIGH
7.4 Jul 14

Memory corruption in NVIDIA TensorRT-LLM allows an attacker with local access to trigger a write-what-where primitive (C

CVE-2026-24229 HIGH
7.3 Jul 14

Missing authentication in NVIDIA TensorRT-LLM for Linux lets an attacker reach the disaggregated orchestrator's FastAPI

CVE-2026-24234 MEDIUM
6.8 Jul 14

Server-side request forgery in NVIDIA TensorRT-LLM for Linux exposes AI inference servers to internal network pivoting v

CVE-2026-24220 MEDIUM
6.4 Jul 14

Unsafe deserialization in NVIDIA TensorRT-LLM's visual gen server through version 1.3.0 rc11 allows a locally privileged

CVE-2026-24259 MEDIUM
6.4 Jul 14

Missing authentication for a critical function in NVIDIA TensorRT-LLM for Linux (all versions through v1.3.0 rc12) allow

Share

EUVD-2026-44470 vulnerability details – vuln.today

This site uses cookies essential for authentication and security. No tracking or analytics cookies are used. Privacy Policy