Raw vector
CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:HSummary
CVE-2026-33298 is a high-severity Heap-based Buffer Overflow (CWE-122) vulnerability in Ggml Llama.Cpp. Its CVSS base score is 7.8 (High).
Operationally, exploitation aligns with the MITRE ATT&CK technique Exploitation for Privilege Escalation (T1068); ranked at the 39th percentile by exploit likelihood (below the median); it is not currently listed in the CISA KEV catalog; a public proof-of-concept is referenced.
This vulnerability is AI-related — categorised as NLP and Transformers; in the Supply Chain and Deployment risk domain.
The strongest mitigations our analysis identified map to SA-11 (Developer Testing and Evaluation) and SI-10 (Information Input Validation) — see the control section below for these in your framework.
Deeper analysis AI-assisted summary
Synthesised by an AI model from the NVD description and linked references — a reading aid, not an authoritative source.
CVE-2026-33298, published on 2026-03-24, is an integer overflow vulnerability (CWE-190) combined with a heap-based buffer overflow (CWE-122) in the `ggml_nbytes` function of llama.cpp, a C/C++ inference engine for large language models (LLMs). Versions prior to b7824 are affected. The flaw allows an attacker to bypass memory validation by crafting a GGUF file with specific tensor dimensions, causing `ggml_nbytes` to return a significantly smaller size than required—such as 4MB instead of exabytes—leading to memory corruption when the tensor is processed. It carries a CVSS v3.1 base score of 7.8 (AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H).
The attack requires local access and user interaction, with no privileges needed. An attacker can supply a malicious GGUF file to a target user running a vulnerable llama.cpp application, tricking them into loading it for LLM inference. This triggers the integer overflow during size calculation, resulting in a heap buffer overflow and potential remote code execution (RCE) through memory corruption.
Mitigation is addressed in the official GitHub security advisory (GHSA-96jg-mvhq-q7q7) and release tag b7824, which contains a fix for the `ggml_nbytes` function. Security practitioners should update to llama.cpp b7824 or later to prevent exploitation.
This vulnerability holds relevance for AI/ML deployments relying on llama.cpp for lightweight, local LLM inference, highlighting risks in file-processing components of such frameworks. No public evidence of real-world exploitation is available.
EU & UK References
- 🇪🇺 ENISA EUVD: EUVD-2026-14668
Vulnerability Data
llama.cpp is an inference of several LLM models in C/C++. Prior to b7824, an integer overflow vulnerability in the `ggml_nbytes` function allows an attacker to bypass memory validation by crafting a GGUF file with specific tensor dimensions. This causes `ggml_nbytes`…
more
to return a significantly smaller size than required (e.g., 4MB instead of Exabytes), leading to a heap-based buffer overflow when the application subsequently processes the tensor. This vulnerability allows potential Remote Code Execution (RCE) via memory corruption. b7824 contains a fix.
- CWE(s)
AI Security AnalysisAI
- AI Category
- NLP and Transformers
- Risk Domain
- Supply Chain and Deployment
- OWASP Top 10 for LLMs 2025
- None mapped
- Classification Reason
- Matched keywords: llama.cpp, llm
Related Threats
MITRE ATT&CK Enterprise Techniques
CVEs Like This One
Affected Assets
Mitigating Controls
Control response
—
—
—
V1.4.1V5.2.6
Mitigating Controls (NIST 800-53 r5) AI
Developer testing and evaluation (including fuzzing and memory-error detectors) can discover heap overflows after they have been coded.
Input validation enforces bounds checking on data written to heap buffers, directly stopping the overflow condition from being introduced.
Security engineering principles require use of memory-safe constructs and bounds-checked allocation routines that avoid introducing heap overflows.
Memory-protection mechanisms limit the ability of a heap overflow to execute attacker-controlled code or corrupt adjacent structures.
Mitigating Controls (NIST CSF 2.0) AI
Derived directly from the weakness types (CWEs) cited in the NVD entry via our AI-authored CWE→CSF cross-walk (authority under review) — links open the control.
Secure-development practices directly require bounds checking and safe memory handling that prevent heap overflows.
Vulnerability scanning and recording can discover heap-overflow flaws but does not prevent their introduction in code.
Timely patching removes known heap-overflow instances after they exist.
Mitigating Controls (ISO/IEC 27001:2022 Annex A) AI
Derived directly from the weakness types (CWEs) cited in the NVD entry via our AI-authored CWE→ISO cross-walk (authority under review) — links open the control.
Security testing in development and acceptance can detect heap overflows before release.
Secure development lifecycle mandates practices that reduce the likelihood of introducing heap overflows.
Application security requirements can specify bounds-checking and safe memory APIs that mitigate heap overflows.
Secure architecture and engineering principles include memory-safety and input-validation controls that address heap overflows.
Secure coding standards directly prescribe techniques (safe functions, bounds checks) that prevent heap-based buffer overflows.
Change management ensures controlled deployment of fixes for discovered heap-overflow vulnerabilities.