Stable Baselines3
Monthly
Unsafe pickle deserialization in the core load APIs of DLR-RM stable-baselines3 up to 2.9.0 lets an attacker who supplies a crafted model, replay buffer, or VecNormalize artifact achieve arbitrary code execution when a victim loads it via PPO.load, load_replay_buffer, or VecNormalize.load. The flaw is a CWE-502 pickle deserialization issue with no safe-mode gate or opt-in guard in these functions, and publicly available exploit code exists; the practical attack path is sharing or publishing a poisoned reinforcement-learning checkpoint that the victim is induced to load. Our assessment scores it CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H - remote delivery is possible but exploitation still hinges on a victim choosing to load the malicious file, and successful loading grants full confidentiality, integrity, and availability impact in the victim's process.
Unsafe pickle deserialization in the core load APIs of DLR-RM stable-baselines3 up to 2.9.0 lets an attacker who supplies a crafted model, replay buffer, or VecNormalize artifact achieve arbitrary code execution when a victim loads it via PPO.load, load_replay_buffer, or VecNormalize.load. The flaw is a CWE-502 pickle deserialization issue with no safe-mode gate or opt-in guard in these functions, and publicly available exploit code exists; the practical attack path is sharing or publishing a poisoned reinforcement-learning checkpoint that the victim is induced to load. Our assessment scores it CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H - remote delivery is possible but exploitation still hinges on a victim choosing to load the malicious file, and successful loading grants full confidentiality, integrity, and availability impact in the victim's process.