Cyber Resilience

CVE-2026-33298

Memory Safety in Ggml Llama.Cpp ≤ b7824

Public PoCMemory Safety
Published
24 March 2026
Modified
30 April 2026
Patch / advisory
CVSS Score v3.1 7.8
Click a component to see what it means
Raw vectorCVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H
EPSS Score 0.0048 39th percentile
Risk Priority 57 floored blend · peak EPSS

Summary

CVE-2026-33298 is a high-severity Heap-based Buffer Overflow (CWE-122) vulnerability in Ggml Llama.Cpp. Its CVSS base score is 7.8 (High).

Operationally, exploitation aligns with the MITRE ATT&CK technique Exploitation for Privilege Escalation (T1068); ranked at the 39th percentile by exploit likelihood (below the median); it is not currently listed in the CISA KEV catalog; a public proof-of-concept is referenced.

This vulnerability is AI-related — categorised as NLP and Transformers; in the Supply Chain and Deployment risk domain.

The strongest mitigations our analysis identified map to SA-11 (Developer Testing and Evaluation) and SI-10 (Information Input Validation) — see the control section below for these in your framework.

Deeper analysis AI-assisted summary

Synthesised by an AI model from the NVD description and linked references — a reading aid, not an authoritative source.

CVE-2026-33298, published on 2026-03-24, is an integer overflow vulnerability (CWE-190) combined with a heap-based buffer overflow (CWE-122) in the `ggml_nbytes` function of llama.cpp, a C/C++ inference engine for large language models (LLMs). Versions prior to b7824 are affected. The flaw allows an attacker to bypass memory validation by crafting a GGUF file with specific tensor dimensions, causing `ggml_nbytes` to return a significantly smaller size than required—such as 4MB instead of exabytes—leading to memory corruption when the tensor is processed. It carries a CVSS v3.1 base score of 7.8 (AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H).

The attack requires local access and user interaction, with no privileges needed. An attacker can supply a malicious GGUF file to a target user running a vulnerable llama.cpp application, tricking them into loading it for LLM inference. This triggers the integer overflow during size calculation, resulting in a heap buffer overflow and potential remote code execution (RCE) through memory corruption.

Mitigation is addressed in the official GitHub security advisory (GHSA-96jg-mvhq-q7q7) and release tag b7824, which contains a fix for the `ggml_nbytes` function. Security practitioners should update to llama.cpp b7824 or later to prevent exploitation.

This vulnerability holds relevance for AI/ML deployments relying on llama.cpp for lightweight, local LLM inference, highlighting risks in file-processing components of such frameworks. No public evidence of real-world exploitation is available.

EU & UK References

Vulnerability Data

llama.cpp is an inference of several LLM models in C/C++. Prior to b7824, an integer overflow vulnerability in the `ggml_nbytes` function allows an attacker to bypass memory validation by crafting a GGUF file with specific tensor dimensions. This causes `ggml_nbytes`…

more

to return a significantly smaller size than required (e.g., 4MB instead of Exabytes), leading to a heap-based buffer overflow when the application subsequently processes the tensor. This vulnerability allows potential Remote Code Execution (RCE) via memory corruption. b7824 contains a fix.

CWE(s)

AI Security AnalysisAI

AI Category
NLP and Transformers
Risk Domain
Supply Chain and Deployment
OWASP Top 10 for LLMs 2025
None mapped
Classification Reason
Matched keywords: llama.cpp, llm

Related Threats

MITRE ATT&CK Enterprise Techniques

T1068 Exploitation for Privilege Escalation Privilege Escalation
Adversaries may exploit software vulnerabilities in an attempt to elevate privileges.
T1190 Exploit Public-Facing Application Initial Access
Adversaries may attempt to exploit a weakness in an Internet-facing host or system to initially access a network.
T1203 Exploitation for Client Execution Execution
Adversaries may exploit software vulnerabilities in client applications to execute code.
T1210 Exploitation of Remote Services Lateral Movement
Adversaries may exploit remote services to gain unauthorized access to internal systems once inside of a network.
T1212 Exploitation for Credential Access Credential Access
Adversaries may exploit software vulnerabilities in an attempt to collect credentials.
Derived from this CVE’s CWE(s) via the direct CWE→ATT&CK cross-walk.

CVEs Like This One

CVE-2024-23496Same product: Ggml Llama.Cpp
CVE-2024-21836Same product: Ggml Llama.Cpp
CVE-2024-21825Same product: Ggml Llama.Cpp
CVE-2024-23605Same product: Ggml Llama.Cpp
CVE-2024-21802Same product: Ggml Llama.Cpp
CVE-2024-42478Same product: Ggml Llama.Cpp
CVE-2025-52566Same product: Ggml Llama.Cpp
CVE-2025-49847Same product: Ggml Llama.Cpp
CVE-2026-34159Same product: Ggml Llama.Cpp
CVE-2024-42479Same product: Ggml Llama.Cpp

Affected Assets

ggml
llama.cpp
≤ b7824

Mitigating Controls

Control response

Prevent
Stop it (NIST 800-53)

Detect
Catch it (NIST detect / respond)

Harden
Shrink the surface (DISA STIG)

Validate
Prove the fix (OWASP ASVS)
  • V1.4.1
  • V5.2.6

Mitigating Controls (NIST 800-53 r5) AI

Developer testing and evaluation (including fuzzing and memory-error detectors) can discover heap overflows after they have been coded.

Input validation enforces bounds checking on data written to heap buffers, directly stopping the overflow condition from being introduced.

Security engineering principles require use of memory-safe constructs and bounds-checked allocation routines that avoid introducing heap overflows.

Memory-protection mechanisms limit the ability of a heap overflow to execute attacker-controlled code or corrupt adjacent structures.

Mitigating Controls (NIST CSF 2.0) AI

Derived directly from the weakness types (CWEs) cited in the NVD entry via our AI-authored CWE→CSF cross-walk (authority under review) — links open the control.

PR.PS-06 full match
prevents

Secure-development practices directly require bounds checking and safe memory handling that prevent heap overflows.

ID.RA-01 partial match
prevents

Vulnerability scanning and recording can discover heap-overflow flaws but does not prevent their introduction in code.

PR.PS-02 partial match
prevents

Timely patching removes known heap-overflow instances after they exist.

Mitigating Controls (ISO/IEC 27001:2022 Annex A) AI

Derived directly from the weakness types (CWEs) cited in the NVD entry via our AI-authored CWE→ISO cross-walk (authority under review) — links open the control.

finds

Security testing in development and acceptance can detect heap overflows before release.

prevents

Secure development lifecycle mandates practices that reduce the likelihood of introducing heap overflows.

prevents

Application security requirements can specify bounds-checking and safe memory APIs that mitigate heap overflows.

prevents

Secure architecture and engineering principles include memory-safety and input-validation controls that address heap overflows.

prevents

Secure coding standards directly prescribe techniques (safe functions, bounds checks) that prevent heap-based buffer overflows.

none

Change management ensures controlled deployment of fixes for discovered heap-overflow vulnerabilities.

References