External risk intelligence

llama.cpp Heap Corruption via Use-After-Free and Double Free in Chat Mapper

CVE advisorySeverity: CRITICAL (CVSS 9.2)

CVE-2026-107183

The vulnerability exists in the llama-server component, which provides a web API interface for interacting with LLMs. As this component is commonly deployed as an internet-facing API service or application backend to facilitate remote access to model completion functions, it is likely to be reachable from the internet in standard deployment patterns.

Use After Free

Halo Surface Signal: 4 out of 5 — likely to be public-facing.

External exposure likelihood

Horizon Alert

Summary of the vulnerability and why it matters

A use-after-free vulnerability in llama.cpp could allow remote attackers to crash the server and potentially gain control of memory. This impacts systems that use this technology for processing chat and completing requests. The primary concern is to confirm if our environment is affected by this specific technical issue.

  • Server crashes possible from remote input.
  • Confirms exposure and relevance of this flaw.
  • Understand technical risk, confirm current exposure.

Attack Path

How an attacker could exploit the issue

An unauthenticated attacker can remotely corrupt heap memory by sending a specially crafted chat parser in a POST request. This crafted request, which includes a tool ID after a tool-close tag, targets a use-after-free and double-free vulnerability within the `common_chat_peg_mapper::map` function. Successful exploitation can lead to a crash of the llama-server and the ability to shape a heap write primitive.

  • Attacker needs network access.
  • Triggered by malformed chat parser input.
  • Risk of server crash and memory corruption.

Live Threat

Current exploitation, exposure, and threat context

This vulnerability could allow unauthenticated remote attackers to crash the `llama-server` service and potentially manipulate its memory. This could occur when a specially crafted chat completion request is sent to the server, triggering a use-after-free and double free vulnerability related to tool mapping.

  • Server stability and heap integrity are at risk.
  • Attackers can send malicious completion requests.
  • Service disruption and memory corruption may occur.

Operational Fix

Recommended remediation, mitigation, and detection steps

This critical vulnerability in llama.cpp's chat parser affects services that expose completion APIs. Infrastructure or platform teams responsible for the llama-server deployment should lead the response. The immediate first step is to confirm the presence and reachability of vulnerable instances, identify the business-criticality, and then engage the accountable owner to prioritize and plan remediation, potentially involving coordination with the vendor or application owners if the server is integrated into a larger product.

  • Own by Infrastructure or Platform Teams.
  • Verify server reachability and business criticality.
  • Plan remediation and vendor coordination.

Supplementary metadata

Validate whether this threat affects your internet-facing exposure.

Halo Threat Intelligence helps prioritize remediation with Halo Surface Signal and H/A/L/O context. Start exposure validation with a free external attack surface trial.

Frequently asked questions

What is llama.cpp and what is it used for?

llama.cpp is an open-source software project that enables the efficient running of Large Language Models (LLMs) on consumer-grade hardware. It provides tools, including the llama-server component, which serves as a backend API. Developers use this server to integrate LLM capabilities, such as chat completion and text generation, into their own applications or custom web interfaces.

How does CVE-2026-107183 impact system memory?

This vulnerability involves a memory management error known as a use-after-free and double free, categorized under CWE-416. Essentially, the software incorrectly manages memory when processing specific chat inputs. By sending a malformed request, an attacker can force the server to reuse or free the same memory space improperly. This corruption can lead to service instability and potentially allow an attacker to influence how the system writes data to its memory heap.

What triggers this vulnerability in the server?

The issue is triggered when the llama-server processes a specifically crafted POST request containing a malformed chat parser input. The attack relies on placing a tool-id tag after a tool-close tag within the request. It is important to note that standard, well-formed chat interactions that follow expected schema rules do not trigger this flaw; the vulnerability specifically exploits the logic error occurring during the improper handling of these out-of-sequence tags.

Is my deployment at risk according to Halo Surface Signal?

Halo Surface Signal indicates that this vulnerability is likely to be reachable from the internet. Because the llama-server component is frequently deployed as an internet-facing API to facilitate remote access for model completion, systems with direct public exposure are at higher risk. If your instance is only accessible within a secure, private internal network, the immediate threat of remote exploitation is significantly lower compared to publicly accessible services.

How should I respond to the CVE-2026-107183 advisory?

Your first step is to perform an inventory of your systems to identify where llama-server is running and whether those instances are reachable via the network. Once located, assess the business criticality of those specific services. Collaborate with your platform or infrastructure teams to plan for updates. Given the nature of this memory corruption flaw, prioritizing the restriction of network access to these servers is a prudent immediate measure until the software is updated.

References