LMDeploy
CVE-2025-67729
HIGH
Severity by source
AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H
Malicious checkpoints are typically network-sourced (AV:N) and need no attacker privileges (PR:N), but the victim must load the file (UI:R); pickle deserialization yields full RCE, so C/I/A all High.
Primary rating from Vendor (github).
CVSS VectorVendor: github
Lifecycle Timeline
2DescriptionCVE.org
LMDeploy is a toolkit for compressing, deploying, and serving LLMs. Prior to version 0.11.1, an insecure deserialization vulnerability exists in lmdeploy where torch.load() is called without the weights_only=True parameter when loading model checkpoint files. This allows an attacker to execute arbitrary code on the victim's machine when they load a malicious .bin or .pt model file. This issue has been patched in version 0.11.1.
AnalysisAI
Arbitrary code execution in InternLM LMDeploy 0.11 and earlier: the toolkit calls torch.load() without weights_only=True when loading model checkpoint files, so a malicious .bin, .pt or .pth checkpoint executes attacker-supplied code during deserialization. Exploitation is unauthenticated in the sense of no credentials (CVSS PR:N) but is gated by user interaction (UI:R) - the victim must deliberately point LMDeploy at an attacker-controlled checkpoint through the quantization/deployment APIs, the TurboMind PytorchLoader, or the VL model loader; requests to LMDeploy's serving path with trusted first-party weights are unaffected, and .safetensors checkpoints use a safe loader. Publicly available exploit code exists in the GitHub advisory, and the issue is resolved in 0.11.1; no confirmed active exploitation (CISA KEV) was identified at time of analysis.
Technical ContextAI
The root cause is CWE-502 (Deserialization of Untrusted Data). PyTorch's torch.load() dispatches to Python's pickle module when reading legacy checkpoint formats, and pickle reconstructs arbitrary objects by invoking callables returned from an object's __reduce__ method - a standard code-execution primitive (os.system, subprocess, etc.). Only the weights_only=True argument restricts the unpickler to tensors and primitive containers, blocking the pickle reduce path. The advisory documents that LMDeploy already used the secure pattern in lmdeploy/pytorch/weight_loader/model_weight_loader.py (line 103) but omitted it in at least six other locations: lmdeploy/vl/model/utils.py (line 22, load_weight_ckpt, which only falls back to torch.load for non-.safetensors files), lmdeploy/turbomind/deploy/loader.py (line 122, PytorchLoader.items), lmdeploy/lite/apis/kv_qparams.py (lines 129-130, key_stats.pth and value_stats.pth), lmdeploy/lite/apis/smooth_quant.py (line 61), lmdeploy/lite/apis/auto_awq.py (line 101, inputs_stats.pth) and lmdeploy/lite/apis/get_small_sharded_hf.py (line 41). The affected technology is the LLM compression/quantization and deployment toolchain that consumes third-party model weights - a supply-chain-relevant surface because users routinely pull checkpoints from model hubs. The upstream fix (commit eb04b428) adds weights_only=True to the auto_awq and get_small_sharded_hf call sites and removes the vulnerable kv_qparams.py module entirely. The RCE/Deserialization/Lmdeploy/Checkpoint tags and the CPE cpe:2.3:a:internlm:lmdeploy:*:*:*:*:*:*:*:* confirm the entire LMDeploy package line is the affected product family. The assessed CVSS vector AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H matches the advisory and the auth/interaction model: no privileges required, but the vector's UI:R reflects that the victim must actively load the hostile checkpoint, so this is not a remote unattended or drive-by trigger.
RemediationAI
Upgrade to LMDeploy 0.11.1 or later - vendor-released patch that applies weights_only=True at the vulnerable torch.load() call sites; the upstream fix is commit https://github.com/InternLM/lmdeploy/commit/eb04b4281c5784a5cff5ea639c8f96b33b3ae5ee and the advisory is https://github.com/InternLM/lmdeploy/security/advisories/GHSA-9pf3-7rrr-x5jh. Where an immediate upgrade is not possible, the highest-value compensating control is to accept only .safetensors checkpoints and reject .bin/.pt/.pth inputs, since LMDeploy's safetensors branch already bypasses pickle; the trade-off is that many community checkpoints are still distributed in legacy PyTorch format, so conversion to safetensors must be done once in a quarantined environment before use. Second, treat every downloaded checkpoint as untrusted code: convert it inside a disposable container or VM with no host credentials, network egress or mounted secrets, accepting the operational cost of a sandboxed conversion step. Third, if you maintain a fork or vendored copy, add weights_only=True explicitly to the affected call sites (vl/model/utils.py load_weight_ckpt, turbomind/deploy/loader.py PytorchLoader.items, lite/apis/smooth_quant.py and lite/apis/auto_awq.py) or backport the upstream commit; note that weights_only=True can fail on checkpoints containing non-tensor Python objects, which may require regenerating the checkpoint with the newer serializer. Do not rely on restricting the lmdeploy serving endpoint, since the trigger is local file loading (UI:R) rather than a network request - access controls on the HTTP API do not reduce this risk. There is no vendor-released workaround beyond the 0.11.1 upgrade.
Remote code execution in InternLM LMDeploy versions 0.9.1 through 0.10.1 allows unauthenticated network attackers to run
Server-Side Request Forgery (SSRF) in InternLM LMDeploy's vision-language module allows remote unauthenticated attackers
Unauthenticated remote code execution in InternLM LMDeploy (versions 0.9.2 through 0.15.x) allows a remote attacker to e
A vulnerability was found in InternLM LMDeploy up to 0.7.1. Rated medium severity (CVSS 4.8), this vulnerability is low
A vulnerability was found in InternLM LMDeploy up to 0.7.1. Rated medium severity (CVSS 4.8), this vulnerability is low
Denial of service in InternLM LMDeploy through 0.17.0 lets unauthenticated remote attackers crash the distributed infere
Memory exhaustion in InternLM LMDeploy through 0.17.0 can be triggered by unauthenticated remote attackers who send repe
Server-side request forgery in InternLM lmdeploy's OpenAI-compatible vision API server lets unauthenticated remote attac
Same weakness CWE-502 – Deserialization of Untrusted Data
View allShare
External POC / Exploit Code
Leaving vuln.today