Skip to main content

llama.cpp CVE-2024-23496

HIGH
Integer Overflow or Wraparound (CWE-190)
2024-02-26 talos-cna@cisco.com
8.8
CVSS 3.1 · NVD
Share

Severity by source

NVD PRIMARY
8.8 HIGH
AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H

Primary rating from NVD · only source for this CVE.

CVSS VectorNVD

CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H
Attack Vector
Network
Attack Complexity
Low
Privileges Required
None
User Interaction
Required
Scope
Unchanged
Confidentiality
High
Integrity
High
Availability
High

DescriptionCVE.org

A heap-based buffer overflow vulnerability exists in the GGUF library gguf_fread_str functionality of llama.cpp Commit 18c2e17. A specially crafted .gguf file can lead to code execution. An attacker can provide a malicious file to trigger this vulnerability.

AnalysisAI

Remote code execution in llama.cpp (commit 18c2e17) occurs when the GGUF library's gguf_fread_str function parses a maliciously crafted .gguf model file, triggering a heap-based buffer overflow rooted in integer overflow handling (CWE-190). Any user or service loading an untrusted GGUF model into a vulnerable llama.cpp build can be compromised, with publicly available exploit code increasing accessibility despite a low EPSS score of 0.15%.

Technical ContextAI

llama.cpp is a widely deployed C/C++ inference engine for LLaMA-family large language models, used in local LLM runners, chat frontends, and embedded AI tooling. The GGUF (GPT-Generated Unified Format) is its binary model serialization format, and gguf_fread_str is the helper that reads length-prefixed strings from .gguf files. CWE-190 (Integer Overflow or Wraparound) indicates that an attacker-controlled length field is processed without sufficient bounds checking, producing an undersized heap allocation followed by an oversized copy - a classic heap buffer overflow primitive that can corrupt adjacent heap metadata or function pointers and yield arbitrary code execution. The affected CPE cpe:2.3:a:ggml:llama.cpp confirms the ggml-maintained upstream project as the impacted component.

RemediationAI

Upstream fix available (PR/commit); released patched version not independently confirmed from the provided data - consult the Talos Intelligence advisory and the ggerganov/llama.cpp GitHub repository for the commit that follows 18c2e17 and rebuild against that or a later release. As a compensating control, only load .gguf files from trusted, signature-verified sources and reject models obtained from untrusted hubs or user uploads, which limits attack surface at the cost of model availability. Where third-party models must be supported, run llama.cpp inside a sandbox (seccomp, container with no network, or a separate low-privilege user) so that successful exploitation does not yield the calling application's privileges, accepting the operational complexity that sandboxing adds. Disabling or pre-screening GGUF file ingestion at upload boundaries (size/length sanity checks) provides partial mitigation but is not a substitute for patching the parser.

CVE-2024-42479 CRITICAL POC
10.0 Aug 12

Arbitrary memory write in llama.cpp's RPC server allows remote unauthenticated attackers to corrupt arbitrary memory add

CVE-2024-21802 HIGH POC
8.8 Feb 26

Remote code execution in llama.cpp (commit 18c2e17) is possible when a user opens a malicious .gguf model file, triggeri

CVE-2024-21825 HIGH POC
8.8 Feb 26

Remote code execution in llama.cpp (commit 18c2e17) is possible when a victim loads a malicious .gguf model file, trigge

CVE-2024-23605 HIGH POC
8.8 Feb 26

Remote code execution in llama.cpp (GGUF library) allows attackers to achieve arbitrary code execution by tricking a use

CVE-2024-21836 HIGH POC
8.8 Feb 26

Heap-based buffer overflow in llama.cpp's GGUF library header parser (commit 18c2e17) enables code execution when a vict

CVE-2026-34159 CRITICAL
9.8 Apr 01

Remote code execution in llama.cpp RPC backend allows unauthenticated attackers with TCP access to achieve arbitrary mem

CVE-2024-42478 MEDIUM POC
5.3 Aug 12

llama.cpp provides LLM inference in C/C++. Rated medium severity (CVSS 5.3), this vulnerability is remotely exploitable,

CVE-2024-32878 HIGH
8.8 Apr 26

Llama.cpp is LLM inference in C/C++. Rated high severity (CVSS 8.8), this vulnerability is remotely exploitable, no auth

CVE-2026-33298 HIGH
7.8 Mar 24

Remote code execution in llama.cpp prior to commit b7824 is possible through a crafted GGUF file that exploits an intege

CVE-2026-27940 HIGH
7.8 Mar 12

Local attackers can achieve heap buffer overflow in llama.cpp versions before b8146 through integer overflow in the GGUF

CVE-2026-17501 MEDIUM
6.9 Jul 27

Remote denial of service in llama.cpp allows unauthenticated attackers to exhaust server resources via crafted JSON sche

CVE-2026-17500 MEDIUM
6.9 Jul 27

Denial of service in ggml-org llama.cpp allows remote attackers to crash the application by sending a crafted JSON schem

Share

CVE-2024-23496 vulnerability details – vuln.today

This site uses cookies essential for authentication and security. No tracking or analytics cookies are used. Privacy Policy