KVM Flaw Breaks Isolation in Nested Virtualization
X-Ops

KVM Flaw Breaks Isolation in Nested Virtualization

KVM Flaw Breaks Isolation in Nested Virtualization

KVM, the Linux kernel hypervisor that powers Proxmox VE, Red Hat Enterprise Linux and most x86-based public clouds, is back in the spotlight following the public disclosure of Zapscape, a use-after-free vulnerability tracked as CVE-2026-64561. The bug lives in the shadow memory management unit, the component KVM uses to translate guest physical addresses, and lets a guest with elevated privileges escape the virtual machine and run code on the host with root permissions. The coordinated disclosure happened on August 6, 2026, after a roughly three-week embargo negotiated with the Linux kernel security team, according to the original Spanish-language coverage by Hispasec's Una al día bulletin and the English-language follow-up by The Hacker News, CloudLinux, TuxCare, Red Hat and the Openwall oss-security mailing list.

Zapscape was found by independent security researcher Hyunwoo Kim, who goes by the handle V4bel and who has emerged over the past months as the most prolific reporter of KVM escape-class flaws. Kim notified the kernel's security contact on July 11, 2026. The CVE was assigned on August 4, the linux-distros pre-notification was sent on August 1, and the upstream patch, the public proof of concept, the GitHub repository with the full research and a detailed technical write-up were all published on August 6. Some outlets, including TuxCare and CloudLinux, date the public release to August 7 because of timezone differences, a minor discrepancy that does not affect the underlying timeline of the bug or the patch.

The root cause is an ordering error in two page fault handlers inside KVM/x86's shadow MMU. When the hypervisor reclaims shadow pages, there is a recursive "zap" path that, under certain conditions, invalidates a guest's page table root. The vulnerable code, introduced in Linux 5.9 by commit f95eec9bed76 on July 8, 2020, validates the shadow root before reclaiming shadow pages rather than after. If the reclaim invalidates the active root, KVM keeps building mappings underneath an already-invalidated root, producing dangling references and post-free write conditions. This is a textbook use-after-free that breaks a core invariant of the hypervisor: invalid pages must never appear on the active MMU page list.

The upstream fix was authored by Sean Christopherson and Paolo Bonzini, two of KVM's long-standing maintainers, and shipped as commit 2abd5287f083, merged on July 21, 2026. The patch reorders the staleness check to run after make_mmu_pages_available() and modifies a dozen lines across arch/x86/kvm/mmu/mmu.c and arch/x86/kvm/mmu/paging_tmpl.h, but closes the use-after-free window completely. Stable kernel releases 6.6.148, 6.12.101, 6.18.42, 7.1.6 and 7.2-rc5 carry the fix, and distributions have been publishing backports throughout August. AlmaLinux shipped its backport as commit 048f1efd0c on August 6, Red Hat issued RHSA-2026:22147 to coordinate enterprise rollouts, and Debian, Ubuntu and Oracle are tracking the patch through their normal security channels.

Real-world exploitation requires three conditions to align, which limits the attack surface but does not eliminate it. First, nested virtualization must be enabled on the host, a setting that has become more common as infrastructure-as-a-service providers expose the ability to run hypervisors inside their virtual machines. Second, the attacker needs root inside the L1 guest, since the code that manipulates the shadow MMU runs in the guest kernel's context. Third, the host CPU must be an AMD SVM processor with NPT enabled, or an Intel Ice Lake-SP or newer system in which case the L1 guest must additionally be exposed to both four-level and five-level EPT page walk lengths. On AMD there is no equivalent EPT length constraint, which keeps the bar slightly lower for AMD-based hosting environments.

When those three conditions are met, the most severe scenario is a full virtual machine escape. A malicious tenant with root inside their L1 guest creates an L2 guest and, through a carefully crafted sequence of page faults, forces the host's shadow MMU to write into freed kernel memory. Execution jumps to attacker-controlled code that runs in ring 0 of the host, with the ability to read, modify or terminate any other virtual machine running on the same hardware, and to pivot across the management network. Kim's proof of concept demonstrates this end-to-end by creating a file named /Zapscape in the host's filesystem owned by root, an unambiguous marker that remote code execution outside the intended isolation boundary has happened.

There is also a second attack vector that does not require nested virtualization at all, and that is what worries shared-hosting providers the most. On CloudLinux 8, 8 LTS, 9, 9 LTS and 10, the /dev/kvm device is world-writable. A user without privileged access, or a compromised website, can create a throwaway virtual machine, enable nested virtualization inside it, and attack the host kernel from within, escalating to root without the provider ever intending to expose nested virtualization to customers. CloudLinux confirmed this scenario in its advisory and TuxCare's own testing observed that the public PoC does not yet produce a complete host escape on CloudLinux-built kernels, instead limited to crashing the L1 guest. That result is not the same as saying the vector is theoretical, only that the demonstration code has not been adapted to the configuration peculiarities of that distribution.

In terms of severity, Red Hat has classified Zapscape as Important with a CVSS 3.1 base score of 7.0 using the vector AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H, reflecting a local attack of high complexity and low privileges. The MITRE-issued CVSS 3.1 score is higher, 8.8, with the vector AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H, where the main difference lies in scope: the impact extends beyond the component being attacked, because escaping the guest necessarily affects the host and all its other workloads. Both scores share the same underlying CWE-825, expired pointer dereference. The discrepancy is a familiar one in the industry: severity depends heavily on the deployment model and on whether the provider exposes nested virtualization to untrusted tenants in the first place.

Until updates land, the community converges on a layered set of temporary mitigations. The fastest, if also the most disruptive, is to unload the kvm_intel and kvm_amd modules with modprobe -r, which removes the host's ability to run virtual machines until the next reboot, and to blacklist them in /etc/modprobe.d/ so they do not auto-load again. For environments that need to keep virtualizing but do not expose nested virtualization, exposure can be limited with echo 'options kvm_intel nested=0' or the equivalent for kvm_amd. On Proxmox VE this requires draining or stopping guests and rebooting the host. Restricting the permissions of /dev/kvm to an administrative group via udev or equivalent rules closes the shared-hosting vector entirely and is a sensible hardening step even after the patch is applied. In every case, the final destination is updating the kernel to a version containing commit 2abd5287f083 or a distribution-specific backport, and rebooting the host.

The patch rollout is well underway. Red Hat's advisory and Bugzilla entry 2510891 are the authoritative reference for enterprise deployments. AlmaLinux, CentOS Stream 10 and the derived rebuilds have the fix in their main branches. TuxCare's KernelCare livepatch service already covers Debian 13 in its main feed, with Ubuntu 24.04 and 22.04 in testing, and is preparing extended-support coverage for 7h, 8, 8 LTS, 9 LTS and Ubuntu 22.04 in its ELS stream. For Ubuntu, the official CVE-2026-64561 page tracks patch status by pocket and architecture. CloudLinux's matrix indicates that kernels for 7h and 8 are still in beta rolling to stable, while 9 and 10 are at the head of the queue, with 7 (kernel 3.10) confirmed not affected because it predates the vulnerable code path.

Zapscape fits into a broader 2026 pattern of KVM escape vulnerabilities, all credited to the same researcher. Two months earlier, on July 6, Januscape (CVE-2026-53359) was disclosed, a sixteen-year-old use-after-free in kvm_mmu_get_child_sp() that had been weaponized as a zero-day in Google's kvmCTF competition before public release, as documented by The Hacker News and the Cloud Security Alliance. Even earlier, ITScape (CVE-2026-46316) targeted the vGIC-ITS emulation in arm64, showing that the hypervisor's attack surface extends well beyond x86. The concentration of discoveries on a single investigator and the short window between disclosures suggest a systematic campaign against KVM's memory translation paths, an area of the kernel that, by its complexity and the strict invariants it must preserve under concurrency, continues to reward careful auditing.

Operationally, the message for KVM-based platform administrators is unambiguous. Even though no major vendor has confirmed active exploitation of CVE-2026-64561 in real environments, the combination of a public PoC, available patches and the qualitative severity of the escape makes this a priority update. Proxmox VE deployments, common in homelabs and at smaller providers, receive the fix through the underlying distribution kernel, typically Debian, Ubuntu or Enterprise Linux, so the upgrade path is the operating system's package manager, not a Proxmox-specific repository. Larger cloud providers must evaluate whether nested virtualization is part of their external offering and, if it is, plan a rolling patch deployment to avoid downtime, leaning on livepatching where available and restricting nested access to tenants that genuinely need it in the interim.

For blue teams watching the attack surface, it is worth remembering that KVM's shadow MMU is not the first line of defense, but the last resort once a guest has already been compromised and is trying to cross the isolation boundary. Long-standing good practices, which Zapscape brings back into focus, include keeping an inventory of nodes with /dev/kvm exposed, applying least privilege to the groups that can access it, and periodically auditing production kernels for known vulnerable versions. Tools like Lynis, OpenSCAP and the maintenance scripts shipped by each distribution can automate parts of this audit, but none of them replaces a disciplined patch management and hypervisor configuration policy. In that sense, Zapscape is a reminder that the complexity of KVM's shadow paging keeps generating high-impact offensive research opportunities, and that defenders need to plan their response with the same rigor attackers bring to discovering new escape routes.

Finally, the human side of the story is worth noting. Hyunwoo Kim, working as V4bel, has gone from a relatively unknown researcher to the sole discoverer of three separate KVM escape-class vulnerabilities in less than three months, two on x86 and one on arm64. Each disclosure has been accompanied by a careful technical write-up, a reproducible repository and a working PoC, which places Kim among the most prolific hypervisor security researchers active today. The KVM community, for its part, has kept up a remarkable response cadence: the Zapscape patch was ready and reviewed ten days after the initial report, and distribution backports followed within days of the embargo lifting. It is a useful reminder that the health of the open source ecosystem is measured not only by code quality but by the quality of the incident response that surrounds it.

For detection engineering, Zapscape also offers a few worthwhile signals to look for in telemetry. Successful exploitation typically requires the attacker to allocate, populate and then reclaim a large number of shadow pages in a short window, which can manifest as a spike in mm::mmu_pages_alloc and a matching drop in mm::mmu_pages_freed from inside a single guest. On the host side, the shadow MMU reclaim path emits messages when it zaps an invalid root, so an aggressive increase in kvm_mmu_zap_oldest_generation traces, especially from a guest that has no business doing high-volume page table work, can be a strong indicator of attempted exploitation. Finally, the public PoC's filesystem side effect, the creation of /Zapscape owned by root, gives defenders a low-noise tripwire that is worth deploying on any KVM host that runs untrusted guests, since legitimate workloads have no reason to write a file with that exact name in the root of the filesystem. None of these signals replace the need to patch, but together they shorten the time between an attempted escape and the moment a security team notices something is wrong.

There is also a broader architectural lesson that the industry is slowly internalizing. Shadow paging in KVM was originally designed as a compatibility layer to support nested virtualization on hardware that did not have efficient nested page table support, and over time it has accumulated corner cases that reflect the long, organic history of the codebase. Hardware vendors have steadily improved their nested virtualization primitives, and most modern CPUs from both AMD and Intel offer hardware-assisted nested page tables that bypass the shadow MMU entirely. That makes the long-term trajectory of KVM clear: shrinking the role of the shadow paging code path in favor of hardware-assisted translation, and using shadow mode only as a fallback for older or niche configurations. Zapscape, Januscape and ITScape are the price the kernel pays for carrying that legacy code forward, and they will keep arriving as long as the shadow MMU remains in production. Defenders planning multi-year capacity for vulnerability management in virtualized Linux should treat the shadow MMU as a recurring source of high-severity findings and budget accordingly, both for patch cycles and for the operational cost of mitigations when patches are not yet available.