Google AX: The New Declarative Orchestration Layer for Autonomous AI Agents That Kubernetes Was Not Built to Handle
Why Kubernetes falls short when the workload is an AI agent
The Google team responsible for production AI projects has published a component that has been rumored for months and has just appeared as an open source release: **AX** (hosted at `agentexecutor.io` and available on GitHub as `google/ax`), a **declarative orchestrator and runtime under Apache 2.0 license** designed specifically to execute and scale autonomous AI agent workloads. AX runs on top of **Agent Substrate**, an infrastructure layer Google has been quietly maturing, and exposes **four Kubernetes-style primitives** (Task, Workspace, Gateway, Model) under the `ax.io/v1alpha1` API group.
This article explains why the launch matters for any organization that is building, evaluating, or simply monitoring the "production AI agents" category, what technical problem AX solves that Kubernetes alone cannot address, and what adoption and architecture decisions platform teams have to make in the coming weeks.
---
The problem Kubernetes does not solve for agents
Kubernetes is, without dispute, the best orchestration abstraction on the market for stateless and batch workloads. But modern AI agents are neither. They are **stateful, bursty, and long-running**.
When an AI agent executes a task, it goes through very distinct phases:
- **Compute-intensive phases**: while it reasons, executes tools, evaluates code, makes calls to external APIs, it consumes GPU and CPU in a sustained way. - **Prolonged waiting phases**: waiting for responses from external language models, waiting for API feedback that can take seconds or minutes, waiting for human intervention for approvals, waiting for environment events. - **Frequent transitions between the two phases**: a single agent can alternate between compute bursts and prolonged waits dozens of times during a single long task.
Kubernetes, by design, keeps pods active throughout their entire lifecycle. This means that during waiting phases, the resources allocated to the agent's sandbox **remain reserved but underutilized**. In a large fleet of agents, this inefficiency multiplies until it becomes a serious financial and operational problem.
On the other hand, if instead of keeping the sandbox active you destroy and recreate it when the agent needs compute again, you introduce **cold starts** — the initialization time of the sandbox — that degrade the interactive experience and break the agent's context.
AX attacks exactly this problem: **task suspension and resumption in less than a second**, preserving agent state while freeing resources during waits.
---
The four primitives: Task, Workspace, Gateway, Model
AX exposes four declarative primitives under the `ax.io/v1alpha1` API group, deliberately inspired by the Kubernetes mental model but designed for agents.
### Task
The **Task** primitive defines the execution lifecycle of an agent task. It specifies:
- Sandbox resource constraints (CPU, memory, GPU). - References to the supporting infrastructure the task needs (workspaces, gateways, models). - Termination criteria and lifecycle hooks.
It is, in a sense, the declarative equivalent of a Kubernetes `Pod`, but with specific semantics for agents: graceful suspension, resumption from the last known state, and handling of external dependencies that may take time.
### Workspace
The **Workspace** primitive handles the assembly of the environment **before** the task begins executing. This is critical because a productive AI agent typically needs:
- Git repositories mounted in the sandbox (so the agent can read and modify code). - **MCP** (Model Context Protocol) servers configured to expose external tools to the agent. - **Skill bundles** — predefined capability packages the agent can invoke. - **Natural language goals** that an initialization agent executes to bootstrap toolchains and system dependencies before the main task starts.
What is interesting is that Workspace allows all of this to be specified declaratively. The operator describes what the agent needs; AX takes care of provisioning it.
### Gateway
The **Gateway** primitive manages the agent's **outbound network security policies**. This is where AX gets serious from a security standpoint:
- Agents can only communicate with **explicitly allowlisted hostnames and ports**. - Credential injection into outbound requests is handled at the Gateway level, not in the agent's code.
This solves one of the most common problems with AI agents in production: the risk of an agent with too much freedom ending up talking to infrastructure it should not. With AX, network policy is declarative and centralized, not implemented as custom `iptables` or proxies in the agent's code.
### Model
The fourth primitive manages references to the AI models the agent will invoke. This includes both internal provider models (Gemini, open models in Vertex AI) and external endpoints compatible with OpenAI or Anthropic. The abstraction allows the agent to be agnostic to model provider, and routing, fallback, and cost tracking decisions to be managed at the orchestrator level.
---
Agent Substrate: the isolation layer
AX does not execute sandboxes directly on the host kernel. It uses **gVisor**, Google's sandbox that intercepts container syscalls and executes them in user space. This provides an additional isolation layer between the agent and the host operating system.
gVisor adds overhead compared to native execution, but the security benefit is significant: an agent that tries to escape its sandbox or execute dangerous syscalls cannot escape to the host kernel without going through the gVisor layer.
The choice of gVisor is not accidental. It is the same technology that powers Google Cloud Run and many Google serverless services, so there is accumulated operational experience and real battle-testing.
---
The known issues Google is being transparent about
It is rare for an open source launch to come with an honest list of known issues. AX does, and that speaks well of the project:
- **Egress proxy dropped connections.** The component mediating agents' outbound connections still has edge cases where legitimate connections drop. This affects the reliability of tasks that depend on frequent external calls. - **Rudimentary secrets management.** Secret management (rotation, injection, scoping) is in an early phase. For organizations with strict compliance requirements, this is probably insufficient without an additional layer on top.
These issues are not dealbreakers for experimentation and non-critical deployments, but they are blockers for production in regulated industries or for scenarios handling sensitive secrets. The transparency about these points suggests Google will address them in upcoming releases, but for now any serious enterprise adoption needs to compensate with additional practices.
---
What AX is not (and why understanding that matters)
AX is deliberately positioned **not as a framework for solo hackers**, but as **a compute primitive for enterprises managing large-scale, long-running agentic fleets**.
This means several things:
- **It is not a replacement for LangChain, LlamaIndex, or CrewAI.** Those are libraries for building individual agents. AX is for executing fleets of those agents in production. - **It does not compete directly with Kubernetes.** It runs on top of Kubernetes (or, at least, on infrastructure compatible with the Kubernetes mental model). What it does is add an agent-specific layer on top. - **It is not a Google Cloud exclusive product.** It is open source under Apache 2.0. It can be deployed on-prem or in any cloud.
For teams that are building individual agents and experimenting, AX is overkill. For teams that are trying to operationalize agents in production, AX is a direct answer to several real problems.
---
Implications for platform teams
If your organization is planning or executing AI agent workloads in production, AX has several concrete implications.
### Reevaluate the orchestration architecture
If you are currently running agents on pure Kubernetes (standard pods, deployments, services), it is worth evaluating whether the inefficiencies AX solves are costing you real money. For large fleets, the answer is usually yes.
### Reevaluate outbound network security policies
AX's Gateway primitive centralizes outbound network policy management. If you currently have policies scattered across Kubernetes NetworkPolicies, sidecar proxies, and corporate firewall rules, AX offers a consolidation point.
### Consider observability
AX exposes agent state in a structured way. This enables observability that is hard to achieve with ad-hoc implementations: which agents are suspended, which tasks are waiting for human input, which models are being invoked, which costs are accumulating.
### Plan integration with the existing stack
AX is not drop-in. It requires thinking about agent runtime architecture from the perspective of declarative primitives, not individual Python scripts. For organizations with existing investments in frameworks like LangChain, there are decisions to make about how both approaches coexist.
---
The strategic question: when to adopt?
Not everyone needs AX today. The honest answer depends on the organization's profile:
- **If you are in the experimental phase with agents** (testing capabilities, evaluating model providers, building proofs of concept), AX is probably too much. Use what you have. - **If you are operationalizing the first agent in production** and that agent is reasonably stable, AX is worth evaluating as a deployment platform, especially if you anticipate scaling to multiple agents. - **If you already have a fleet of agents in production** and are suffering the problems AX solves (resource underutilization during waits, cold starts degrading UX, manual per-agent network policy management), AX is a direct answer worth trying.
What does not make sense is waiting a year to see how it matures. The agent orchestration space is moving fast, and the platform decision you make in the next 6 months will be hard to reverse later.
---
Conclusion: a primitive that was missing from the stack
AX is not just another project in the AI space. It is the explicit recognition that **AI agents are a distinct workload category** that deserves specific orchestration primitives, not forced adaptations of primitives designed for microservices or batch jobs.
Google is betting that this category will dominate the next decade of cloud infrastructure, and is making the code available so the industry can build on that foundation. That is significant, regardless of whether AX ends up being the standard or simply one of the contenders.
For platform teams, the clear signal is: **start thinking about your agents as a workload category of their own**. AX is one of the first serious materializations of that category, and even if you do not adopt it, the mental model it proposes (task suspension, gateway policy, declarative workspace) will influence how alternatives are built in the coming years.
---
*Sources: Google's official announcement of AX; documentation at `agentexecutor.io`; GitHub repository `google/ax`; technical analysis by Olimpiu Pop published on InfoQ; gVisor documentation on container sandboxing.*