Unikraft and the AI infrastructure scale problem: stuffing a million sandboxes into a single server
Executive summary
At a recent talk in London, Felipe Huici, CEO and co-founder of Unikraft, asked his audience a direct question: how many virtual machines can fit in a 48-core server? The options he offered were tongue-in-cheek: one full Ubuntu, tens, hundreds if the operator feels adventurous, or, if you can scale them to zero when idle, an even higher density. The talk was titled "Fixing the AI Infra Scale Problem by Stuffing 1M Sandboxes in a Single Server" and its central thesis is that the bottleneck of the new generation of AI infrastructure is not compute but isolation density. The sandboxes that agentic AI needs to execute no longer fit in the container model, nor even in the classical microVM model. What is needed is a redesign of the isolation primitive from scratch, and Unikraft has built its entire platform around that idea. This article explains what Unikraft is, why isolation matters more than ever for AI workloads, what its relation is to Firecracker and to unikernels, and what architectural decisions you can make tomorrow if you are building agentic code execution platforms.
The changing meaning of the word "sandbox"
For nearly a decade, sandbox was synonymous with container. Docker popularized the term in 2013 and, since then, most developers associate sandbox with container image, container runtime, and Kubernetes orchestration. Felipe Huici opens his talk by admitting that the term has been recycled: today, when people talk about sandboxes in the AI context, they almost never mean containers. They mean isolated environments where an AI agent — whether an LLM capable of executing code, a multi-agent system orchestrating tools, or a pipeline evaluating generated code — can run instructions safely, without affecting the rest of the system.
This semantic shift is not cosmetic. It is structural. A Docker container shares the host's kernel. An AI agent executing arbitrary code in a container is, technically, executing that code with access to the host kernel. The standard mitigations — seccomp, AppArmor, namespaces — reduce the attack surface but do not eliminate it. A container escape means compromising the host. And an AI agent is, by definition, a piece of software that executes code that was not audited by a human before running. If your agent executes a thousand instructions per minute, the cumulative probability that one of them is malicious or simply defective enough to enable an escape is non-zero.
That is why the isolation primitive matters. A Firecracker-based microVM no longer shares a kernel with the host: each VM has its own kernel, its own root filesystem, and its own set of virtualized devices. A VM escape still has to traverse the hypervisor before touching the host, and the hypervisor is a much smaller code base than a full Linux kernel. That is the operational difference between "I trust my sandboxes will not be escaped" and "my architecture is designed assuming escape attempts will happen".
What Unikraft is and what it brings to the debate
Unikraft is not a new hypervisor. It is a platform for building cloud platforms. The company commercializes the idea that any cloud platform — from an edge function runtime to an agentic execution service — should be built on a foundation that offers native isolation density, without sacrificing cold boot speed. The piece Unikraft brings to the ecosystem is a framework for composing unikernels and microVMs optimized for specific workloads, along with tooling to deploy them in production.
Huici's talk walks through the isolation primitives available: standard virtual machines, microVMs, unikernels, containers, and language-level runtime isolation. The point he arrives at is that each has its place, but containers alone are not sufficient for workloads where the executed code is generated by an AI model. MicroVMs are a step forward, but a classic microVM still loads a full Linux kernel and a reduced userland. A unikernel goes further: instead of running a minimal Linux distribution inside the VM, the unikernel compiles the application together with a minimal kernel that only contains the drivers and subsystems that application needs. The result is a much smaller image, with a TCB measured in hundreds of thousands of lines of code instead of tens of millions.
Cold boot is where the primitive shines. Huici cites measurements where Unikraft-optimized microVMs boot in a few milliseconds. For agentic workloads, where every tool call may require a fresh sandbox, milliseconds versus seconds is the difference between a fluid user experience and a system that feels slow. For platforms that execute millions of sandboxes per hour, milliseconds versus seconds is the difference between a viable infrastructure cost and one that breaks the business model.
The TCB argument
The trusted computing base, or TCB, is the set of software that must be correct for the system to function securely. If your TCB is small, you have less code to audit, less attack surface, and lower probability that a vulnerability slips through. If your TCB is large, the opposite holds.
For a standard VM, the TCB includes the hypervisor, the VMM, and, inside the VM, the full Linux kernel. The Linux kernel, in its modern versions, runs around 30 million lines of code. Every release includes thousands of patches; every subsystem — network, storage, scheduler, drivers — is a potential vector. If you trust your kernel to be free of vulnerabilities, you are trusting that 30 million lines contain no exploitable errors. It is a bet that history has shown to be a losing one.
For a container, the TCB is worse: it includes the host's full kernel, plus the container runtime, plus the kernel's namespaces and cgroups. The container does not isolate from the kernel; it shares the kernel. That means a vulnerability in any kernel subsystem is a vulnerability in every container running on that host. Seccomp and AppArmor reduce the risk by limiting the available syscalls, but they do not remove the kernel from the TCB.
For a unikernel, the TCB is drastically smaller: the minimal kernel compiled together with the application, plus the hypervisor. If the hypervisor is KVM or Firecracker, we are talking about a total TCB of hundreds of thousands of lines of code, possibly less if only the necessary drivers are compiled. This does not guarantee absolute security, but it radically changes the audit model: instead of auditing 30 million lines, you audit a few tens of thousands, and your attack surface fits in a spreadsheet.
Why isolation matters for agentic AI
Agentic AI is the use case that has put sandboxing back at the center of the infrastructure debate. An agent is, in essence, an LLM that decides which tools to call, executes those tools, observes the results, and iterates. Typical tools include Python code execution, HTTP calls to external APIs, file read and write, and shells. Each of these tools is, from a security perspective, a potential attack vector.
Three concrete risks make isolation matter more than ever:
The first is the prompt injection risk. An agent that summarizes a web page is exposing the LLM's context to the content of that page. An attacker placing malicious instructions on a web page can manipulate the agent. Without robust isolation between the agent and the system, those instructions can end up executing on the host with undue privileges.
The second is the unaudited code risk. An agent that generates Python code and executes it is running code that no human reviewed. The error rate in LLM-generated code remains high enough that a non-trivial percentage of executions fail, and a smaller but non-negligible percentage produces unintended side effects: deleting files outside the expected directory, opening sockets to arbitrary destinations, consuming memory until the host is affected.
The third is the cost risk. An agent in a loop, iterating without progress, can consume host resources to the point of degrading the service. Without isolation, this consumption can affect other tenants. With per-VM isolation, the worst case is that specific VM becomes unusable; the rest of the system keeps working.
For platforms running agents at scale, the combination of these three risks makes container-based isolation insufficient. The leading platforms in the space — Anthropic with its execution environment, OpenAI with its tool runtime, the infrastructure of several startups — are migrating toward microVM- or unikernel-based architectures precisely because the container model, while useful in many contexts, does not deliver the guarantees agentic AI demands.
Density as a business problem
The number Huici puts on the table — a million sandboxes in one server — is not a vanity benchmark. It is a bound that defines which business models are viable. If your platform executes one sandbox per tool call of an agent, and an active agent generates tens of calls per minute, and you have millions of active agents, then the unit cost per sandbox determines your margin. If your sandbox consumes 512 MB of RAM, a server with 512 GB of RAM runs a thousand sandboxes. If your sandbox consumes 50 MB, the same server runs ten thousand. If your sandbox consumes 5 MB, a hundred thousand. And if your sandbox consumes 0.5 MB, a million.
Memory cost is not the only factor. Cold boot cost matters too: if every sandbox takes 5 seconds to start, your platform can only create 12 sandboxes per second per node, and you need hundreds of nodes to sustain a traffic of thousands of sandboxes per second. If cold boot drops to 50 milliseconds, a single node can create 20 sandboxes per second, and the economics change radically.
The third factor is operational cost. Maintaining an optimized unikernel or microVM requires expertise that maintaining a Docker container does not. Unikraft, precisely, offers that expertise as a product: instead of each platform team having to learn how to compile unikernels, it can use Unikraft's framework to build images optimized for its specific workloads.
When unikernel, when microVM, when container
Not every workload needs unikernels. The right decision depends on three axes: isolation requirements, density requirements, and portability requirements.
If your workload is a traditional stateless web service receiving HTTP traffic, processing, and responding, a container is probably enough. Kernel isolation is acceptable because the code is audited, traffic is predictable, and the attack surface is limited. The benefits of unikernel — smaller TCB, lower memory — do not compensate for the loss of familiarity and tooling.
If your workload executes AI-generated code or unaudited third-party code, step up to microVM. Firecracker is the most popular option; VMs boot in hundreds of milliseconds, the TCB is reasonable, and Linux compatibility is high. Memory cost is higher than unikernel, but the operational model is familiar.
If your workload is a high-frequency AI agent that needs to scale to millions of concurrent sandboxes, unikernel starts to make sense. The smaller TCB reduces security risk, the smaller memory footprint increases density, and the faster cold boot improves latency. The price is loss of portability: a unikernel compiled for KVM does not run outside KVM without recompilation.
If your workload executes code that needs access to specific drivers, specialized hardware, or kernel extensions, unikernel can be difficult. Some drivers are not available in unikernel-compatible form. In these cases, microVM with standard Linux kernel may be the best choice.
The cost of choosing wrong
The most expensive mistake a platform team can make today is underestimating the shift in requirements that agentic AI imposes. Build a container-based platform to run AI agents, discover the isolation is insufficient, migrate to microVMs, discover the density does not scale, migrate to unikernels, and only then realize the shortest path would have been unikernels from the start — that cycle costs months of engineering work and hundreds of thousands of dollars in wasted infrastructure.
The opposite mistake is also common: invest in unikernels for workloads that do not need them, pay the complexity and tooling cost, and discover that the benefits do not materialize because the workload was too simple. In this case the cost is less dramatic but still real: engineering time spent on an unnecessary optimization.
The right decision depends on asking the right questions at the start: what kind of code will the sandbox execute? Who wrote it? What happens if the sandbox is compromised? How many sandboxes do I need to run concurrently? What is the acceptable latency for creating a fresh sandbox? If the answers to those questions include "AI-generated code", "many", and "low latency", unikernels or microVMs are probably the right answer.
The future of the space
The sandboxing space for AI is evolving rapidly. Several startups are building platforms optimized for agentic workloads. Hyperscale cloud providers are launching secure execution services based on microVMs. The open source community is producing tooling to compile unikernels without deep expertise.
The question is not whether VM-based sandboxing will replace containers for AI, but when and for which workloads. Coexistence is likely: containers for traditional services, microVMs for general agentic workloads, unikernels for high-density or high-security-sensitivity agentic workloads. The right choice depends on context.
What is clear is that the isolation primitive is an architecture decision that must be made at the start of an AI platform design, not at the end. Changing the primitive after the system is in production is costly and risky. Teams building platforms for agentic AI should take the time to evaluate unikernels and microVMs as part of their initial design, not as a late optimization.
Conclusion
Unikraft and the hyperdense VM-based isolation primitive are not a passing fad. They are the right answer to a real problem: agentic AI runs code that was not audited by humans, at scales the container model cannot support, with security requirements that the shared kernel cannot guarantee. For teams building agentic execution platforms today, optimized unikernels and microVMs represent the frontier of the state of the art.
If you are designing an AI platform and have not seriously evaluated unikernels or optimized microVMs, this is the moment. Felipe Huici's talk is an excellent starting point; the Unikraft framework and similar projects give you the tooling to begin. Density, cold boot, and TCB are not late optimizations: they are architecture requirements from day one.
The cost of doing it right from the start is a fraction of the cost of migrating later. And the migration, in this space, is not optional: when your agentic workloads grow to the point of needing hyperdense sandboxes, containers will fall short. Better to discover it now than in production.