Severity by source
AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H
Network-reachable management endpoint is unauthenticated by default (PR:N) and the pickle fallback yields deterministic RCE (AC:L), giving full C/I/A impact on the host.
Primary rating from Vendor (certcc).
CVSS VectorVendor: certcc
Lifecycle Timeline
4DescriptionCVE.org
SGLang contains a RCE vulnerability when attempting to load model weights from a HuggingFace repository, specifically within the /update_weights_from_disk, where torch.load(..., weights_only=False) fallback enables pickle deserialization of .bin files.
AnalysisAI
Remote code execution in SGLang (versions up to and including 0.5.15) allows attackers to run arbitrary code on the inference server by abusing the /update_weights_from_disk endpoint, which falls back to torch.load(..., weights_only=False) and thus deserializes attacker-controlled pickle streams embedded in .bin model-weight files. Because SGLang's HTTP serving API is typically exposed without authentication, an attacker able to reach the endpoint and influence the loaded weights path can achieve code execution as the serving process. This is a CWE-502 deserialization flaw rated CVSS 9.8; a vendor security advisory (GHSA-wf98-gv64-5wrf) and a public technical disclosure exist, though EPSS remains low (0.26%, 18th percentile) and it is not in CISA KEV - no public exploit identified at time of analysis beyond the disclosure write-up.
Technical ContextAI
SGLang is a high-performance serving and runtime framework for large language models. The vulnerability lives in its weight-management control plane: the /update_weights_from_disk HTTP endpoint lets an operator hot-swap model weights from a filesystem path, often sourced from a HuggingFace repository. When loading legacy PyTorch .bin checkpoints, the code path falls back to torch.load with weights_only=False. PyTorch's default serialization format is Python pickle, and weights_only=False disables the safe-tensor-only guard, so torch.load will execute any __reduce__ / __setstate__ gadget encoded in the pickle stream during deserialization. This is the classic CWE-502 (Deserialization of Untrusted Data) root cause: the .bin file is not merely parsed as tensor data but is treated as executable object state. The affected component is cpe:2.3:a:sglang:sglang, all versions from 0 through 0.5.15 per the ENISA EUVD record.
RemediationAI
Upgrade SGLang to a release newer than 0.5.15 that addresses GHSA-wf98-gv64-5wrf (consult the advisory at https://github.com/sgl-project/sglang/security/advisories/GHSA-wf98-gv64-5wrf for the exact fixed version; a released patched version is not independently confirmed from the provided data, so verify the fixed tag in the advisory before deploying). Until patched, apply compensating controls: do not expose the SGLang HTTP API to untrusted networks - bind it to localhost or place it behind an authenticated reverse proxy and restrict the /update_weights_from_disk route specifically, accepting that this disables remote hot-swapping of weights; load only .safetensors weights and avoid legacy .bin checkpoints so the pickle path is never taken (trade-off: some HuggingFace repos ship only .bin files and would need conversion); and constrain the weights source to a vetted local directory with strict filesystem permissions so an attacker cannot substitute a malicious .bin. Network-segment inference hosts and monitor for unexpected child processes spawned by the serving process as a detection backstop.
Remote code execution in SGLang (versions up to and including 0.5.15) allows unauthenticated attackers to run arbitrary
SGLang's multimodal generation module deserializes untrusted data with pickle.loads() over an unauthenticated ZMQ broker
SGLang's encoder parallel disaggregation system is vulnerable to unauthenticated RCE through pickle deserialization in t
Unauthenticated remote code execution in SGLang (the LLM/multimodal generation serving runtime) affecting version 5.10 a
Remote code execution in SGLang 0.5.9's /v1/rerank endpoint allows unauthenticated attackers to execute arbitrary code b
Remote code execution in SGLang AI inference servers allows unauthenticated attackers to run arbitrary code through the
Remote code execution in SGLang (versions ≤ v0.5.15) allows attackers to achieve arbitrary code execution through the op
Unauthenticated remote code execution affects SGLang, an LLM/multimodal inference-serving framework, at version 5.10, wh
Unauthenticated remote code execution in SGLang versions through v0.5.20 arises from unsafe pickle deserialization in th
Unauthenticated remote code execution in SGLang 0.5.11 through 0.5.14 occurs when the multimodal generation runtime is l
Unauthenticated remote code execution in SGLang (versions 0 through 0.5.14) arises when the expert-parallel backup subsy
Arbitrary file write in SGLang's multimodal generation runtime (version 5.10) allows a remote, unauthenticated attacker
Same weakness CWE-502 – Deserialization of Untrusted Data
View allSame technique Deserialization
View allShare
External POC / Exploit Code
Leaving vuln.today
EUVD-2026-51264
GHSA-r344-357p-w9pp