Severity by source
AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N
Network-accessible endpoints require no authentication when API keys are absent; full confidentiality impact as all model weights are exposed; no integrity or availability impact applies.
Primary rating from Vendor (certcc).
CVSS VectorVendor: certcc
Lifecycle Timeline
4DescriptionCVE.org
SGLang contains a model weight exfiltration vulnerability when no API keys are configured, as SGLang will expose two endpoints that allow a remote attacker to trigger distributed weight broadcasting using NCCL and then triggering data transfer, attackers can exfiltrate all model weights.
AnalysisAI
SGLang versions up to and including v0.5.15 expose unauthenticated network endpoints that allow complete model weight exfiltration when the server is deployed without API key authentication. Attackers can abuse distributed weight broadcasting via NCCL - the same mechanism used for legitimate inter-GPU weight distribution - to redirect all model weights to an attacker-controlled destination, resulting in total confidentiality loss of what may be highly valuable proprietary ML assets. No public exploit code or CISA KEV listing is confirmed at time of analysis; EPSS at 0.19% (9th percentile) reflects low observed exploitation activity, consistent with this being a targeted threat against specialized ML infrastructure rather than opportunistic mass exploitation.
Technical ContextAI
SGLang (cpe:2.3:a:sglang:sglang:*:*:*:*:*:*:*:*) is an ML inference serving framework that supports distributed inference across GPU clusters using NVIDIA Collective Communications Library (NCCL), a high-throughput library designed for collective tensor and weight broadcasting operations. The vulnerability is rooted in CWE-306 (Missing Authentication for Critical Function): two API endpoints controlling NCCL-based distributed weight broadcasting and data transfer initiation are exposed over the network without any authentication gate when the server is started without API key configuration. These endpoints appear designed for internal orchestration within trusted distributed inference clusters but are reachable by any network-adjacent host in the absence of application-layer authentication, making the distributed weight transfer machinery directly accessible to unauthenticated remote actors.
RemediationAI
Upgrade SGLang to a version after v0.5.15, as directed by the upstream security advisory at https://github.com/sgl-project/sglang/security/advisories/GHSA-cpqq-22v3-2wfm; the exact first patched release version is not independently confirmed in available data and should be verified against the official advisory before deployment. As an immediate compensating control, enable API key authentication for all SGLang server instances - this single configuration change eliminates the attack surface described in this CVE without requiring a version upgrade, though upgrading remains the recommended long-term fix. Additionally, restrict network access to SGLang API endpoints via firewall rules or network segmentation so that only trusted orchestration services can reach the distributed inference endpoints; this defense-in-depth measure reduces exposure even if authentication is not configured, but does not address the underlying CWE-306 flaw and should not substitute for proper authentication. Avoid exposing SGLang inference servers directly to public or untrusted networks under any configuration.
Remote code execution in SGLang (versions up to and including 0.5.15) allows unauthenticated attackers to run arbitrary
SGLang's multimodal generation module deserializes untrusted data with pickle.loads() over an unauthenticated ZMQ broker
SGLang's encoder parallel disaggregation system is vulnerable to unauthenticated RCE through pickle deserialization in t
Unauthenticated remote code execution in SGLang (the LLM/multimodal generation serving runtime) affecting version 5.10 a
Remote code execution in SGLang 0.5.9's /v1/rerank endpoint allows unauthenticated attackers to execute arbitrary code b
Remote code execution in SGLang (versions up to and including 0.5.15) allows attackers to run arbitrary code on the infe
Remote code execution in SGLang AI inference servers allows unauthenticated attackers to run arbitrary code through the
Remote code execution in SGLang (versions ≤ v0.5.15) allows attackers to achieve arbitrary code execution through the op
Unauthenticated remote code execution affects SGLang, an LLM/multimodal inference-serving framework, at version 5.10, wh
Unauthenticated remote code execution in SGLang versions through v0.5.20 arises from unsafe pickle deserialization in th
Unauthenticated remote code execution in SGLang 0.5.11 through 0.5.14 occurs when the multimodal generation runtime is l
Unauthenticated remote code execution in SGLang (versions 0 through 0.5.14) arises when the expert-parallel backup subsy
Same technique Authentication Bypass
View allShare
External POC / Exploit Code
Leaving vuln.today
EUVD-2026-51266
GHSA-x6x2-r266-cx9p