Zapscape (CVE-2026-64561): how a privileged L1 guest can escape KVM to the host
# Zapscape (CVE-2026-64561): how a privileged L1 guest can escape KVM to the host
When a guest with root privileges inside a virtual machine manages to execute code in the hypervisor that contains it, the security model of virtualization breaks at its foundation. The promise that cloud providers and data center administrators make to their customers — isolation between tenants, separation between workloads, guarantee that a compromised VM does not affect the others — depends on that boundary not being crossed. CVE-2026-64561, publicly disclosed as Zapscape on August 6, 2026, describes exactly that crossing: a use-after-free flaw in the shadow MMU emulation of KVM/x86 that allows a privileged guest to escape to the host and execute code with kernel privileges on the host system. The vulnerability is the work of researcher Hyunwoo Kim, who had been publishing a series of similar findings in the Linux virtualization stack — Zapscape is the third in a trilogy after ITScape (CVE-2026-46316, on KVM/arm64) and Januscape (CVE-2026-53359, also on KVM/x86 shadow-MMU).
What the shadow MMU is and why it matters here
KVM, the Linux kernel-based hypervisor, offers hardware-assisted virtualization acceleration through CPU extensions like Intel VT-x and AMD-V. The hardware-assisted MMU (EPT on Intel, NPT on AMD) translates guest addresses to host addresses directly, without the hypervisor intervening on each translation. But when a guest runs another hypervisor — that is, when there is nested virtualization, where an L0 guest hosts an L1 guest that in turn might host an L2 guest — address translation becomes recursive. KVM/x86 handles that recursion using a software component called the shadow MMU, which maintains the so-called shadow page tables: tables that translate L2 guest addresses to L1 guest addresses, which will then be translated by hardware to the host.
The shadow MMU is a complex piece. It handles pages, page table roots, translation caches, and reference counting. Every time a shadow MMU page is no longer needed, the code frees it; every time it is needed again, it is reclaimed. The `mmu_page_zap_pte()` function, which Zapscape puts at the center of the vulnerability, handles erasing entries in shadow page tables during that reclamation. The problem is that this function does not adequately verify a reference counter (`root_count`) that should protect it against race conditions in which the page table root has been invalidated while the function is still operating on it.
How the exploit works
The exploit leverages an incorrect ordering between two operations: the check of whether a shadow MMU root is stale, and the execution of an operation that can invalidate that root. In the vulnerable code, KVM first checks whether the root has been invalidated by MMU page reclamation, then proceeds to use that root to map or fetch more pages. But the reclamation can happen between the check and the use, invalidating the root after KVM has already decided it was valid. KVM then continues operating with an invalid root — a pointer pointing to memory that no longer belongs to it — and that is literally a use-after-free.
In terms of impact, a use-after-free on Linux kernel structures that control shadow page tables ends, after several iterations, giving the attacker control over what code executes in the host kernel context. Researcher Hyunwoo Kim demonstrated an exploit chain that ends running commands on the host with root privileges. The PoC code is public and available in the researcher's repository.
The preconditions: what an attacker needs
Zapscape is not a flaw that any guest can exploit. The main precondition is that the attacker already has kernel privileges inside the L1 guest — that is, is already root inside the VM. That reduces the immediate attack surface to scenarios where the attacker already has some level of compromise: a multi-tenant VM where a customer has managed to escalate privileges, an internal guest that the attacker controls through other paths, or a nested virtualization environment where L1 is accessible from L2.
On Intel, there is an additional condition: both EPT page-walk length 4 and page-walk length 5 must be exposed to L1. This matters because nested virtualization on Intel normally limits the page-walk lengths available to L1, and only when both lengths are exposed does the requirement hold. On AMD, that condition does not exist: exploitation is viable without specific page-walk configurations.
The result is that Zapscape is especially severe in nested virtualization scenarios with untrusted guests. Public cloud providers supporting nested virtualization, research environments running hypervisors inside VMs, and organizations running multi-tenant workloads over KVM should treat this CVE with maximum seriousness.
The disclosure timeline
Zapscape's disclosure followed a reasonably orderly process. Kim reported the flaw to security@kernel.org on July 11, 2026. A patch was developed and merged into the kernel on July 21, with commit `2abd5287f083`. On August 1, the issue was submitted to the linux-distros list under a five-day embargo. CVE-2026-64561 was assigned on August 4. Public disclosure followed on August 6.
The fix, according to the commit, moves the stale-root check after the call to `make_mmu_pages_available()`. If reclamation invalidates the current root, KVM now restarts the fault with `RET_PF_RETRY` instead of continuing to map or fetch with the invalid root. It is a surgical correction that attacks exactly the race condition without touching the general logic of the shadow MMU.
Immediate mitigations for those who cannot patch today
For organizations running KVM hosts with nested virtualization exposed to untrusted guests, there are three mitigations available before being able to apply the patch.
The first is disabling nested virtualization. On Intel systems, where the attack requires nested EPT page-walk with both lengths exposed, eliminating nested virtualization closes the primary vector. On AMD systems, where the vulnerability can also be triggered without nested virtualization, disabling it reduces but does not eliminate exposure. The specific configuration varies by distribution: in many kernels, `modprobe -r kvm-intel nested=0` or its equivalent in `/etc/modprobe.d/` takes effect. It is a mitigation that has functional cost — legitimate users who need nested virt lose that capability — but in scenarios where nested virt is not a business feature, it is the cleanest option.
The second is updating the kernel. The major Linux vendors have already published kernels with the backport of the fix. For Red Hat Enterprise Linux and derivatives like Rocky Linux, versions 8, 9, and 10 have updates; CIQ published a specific mitigation guide for those variants. For Ubuntu, the updated HWE and GA kernels include the backport. For Debian, the kernels in stable have the fix. For those who compile their own kernel, applying commit `2abd5287f083` or a later one closes the vulnerability.
The third is reviewing who has access to L1. If your virtualization environment does not need nested virtualization, disable it by default and enable it only in specific tenants or workloads where it is strictly necessary. If you need it, segment: do not allow L1 guests to run arbitrary workloads or be controlled by external users without an additional level of review. Exploitation requires L1 root, so any measure that raises the cost of obtaining root within L1 partially mitigates the risk.
Detection and monitoring
Detecting a Zapscape exploitation attempt is not trivial. The exploit operates inside the L1 guest and symptoms appear on the host as anomalous kernel behavior. Signals to look for include kernel crashes in shadow MMU code — specifically in functions like `mmu_page_zap_pte()` or in MMU fault handlers — that appear shortly after an L1 guest has performed intensive page table management operations. A kernel panic with a stack trace pointing to those functions, on a host running nested virtualization, is a strong candidate for investigation.
It is also worth monitoring QEMU and libvirt logs for unexpected KVM operations. If your host runs QEMU with monitor enabled, reviewing the monitor for anomalous calls to `dump-guest-memory` or for changes in the guest's page table configuration can give clues. But the most reliable detection remains prevention: patch, restrict nested virt, do not expose untrusted guests to L1.
Kim's trilogy and what it suggests about the surface
Zapscape is not an isolated case. It is the third of three VM escape vulnerabilities published by Hyunwoo Kim in 2026: ITScape on KVM/arm64 in June, Januscape on KVM/x86 shadow MMU in July, Zapscape on KVM/x86 shadow MMU in August. Three findings in three months, all against the same virtualization stack, all allowing escape from guest to host. That is not bad luck; it is a surface that admits more investigation.
For platform teams running KVM virtualization in production, the implication is that they must assume more CVEs of this class will come. The reasonable response is structural: treat the virtualization stack as critical security code, not as commodity infrastructure. That means investing in specific monitoring, in accelerated patch processes for kernels, in strict segmentation between tenants, and in a threat modeling program that explicitly considers the "attacker with root inside guest" case. If those elements are in place, the next Zapscape variant will face an organization that responds in hours, not weeks.
Why cloud providers should worry especially
Public cloud providers offering instances with nested virtualization support are the worst-case scenario for Zapscape. A normal cloud instance — without nested virt — runs directly on the provider's hardware, with KVM as L0 and the customer's operating system as L1. That is already vulnerable to any L1→L0 escape if it existed; but the provider's business model depends on it not existing. Zapscape requires that L1 is already running a hypervisor that in turn handles an L2; that is, it requires nested virt enabled and an attacker-controlled L2. That is a rare but not impossible scenario, especially in providers that offer "bare metal as a service" where the customer controls L1 completely, or in providers that have decided to enable nested virt by default to support legitimate use cases.
For a cloud provider, the correct mitigation is not just patching — that is done by default — but reconsidering when and how nested virtualization is offered. Enabling it by default for all tenants unnecessarily expands the attack surface. Enabling it only upon explicit request, with review of the justification and with reinforced monitoring over those specific tenants, reduces the surface without losing functional capability. That decision is a product and platform decision, not just a security one.
What an attacker with root in L1 can really do after
Once the Zapscape exploitation succeeds and the attacker executes code on the host with kernel privileges, the options are the same as an attacker that managed root on the host directly. They can read memory from any other VM on the same host. They can modify hypervisor behavior to intercept operations of other guests. They can plant persistence on the host — malicious kernel modules, modified iptables rules, altered systemd services — that survive host reboots or even individual guest reinstalls. They can pivot to the provider's management network, where provisioning and monitoring tools live. In short: a successful L1→L0 escape is equivalent to compromising the physical host, with everything that implies in terms of access to other tenants, to the management infrastructure, and to the secrets residing on that host.
That is why the right response to Zapscape is not only to patch; it is to review the security architecture of the virtualization layer as a whole. The useful question a security team can ask is: if an attacker managed root on one of our KVM hosts today, what data and systems would remain accessible? The answer to that question reveals the potential consequences of Zapscape and of any future variant of the same class.
Closing
Zapscape is a reminder that virtualization, although it has been maturing for decades, is not an infallible security layer. The combination of an operation-ordering flaw in complex code, a realistic precondition in multi-tenant environments, and a publicly available PoC, makes it a serious vulnerability for any organization running KVM with exposed nested virtualization. The immediate action is to patch; the structural action is to review segmentation and L1 access. And the conversation to keep open is the one that has been going on since ITScape: the Linux virtualization stack deserves the same level of security investment as any other piece of critical infrastructure, because the consequences of a flaw there are catastrophic for the entire multi-tenant model.