Observability and Cost Controls for AI Agents: Stop Your Token Bill From Becoming an Incident
# Observability and Cost Controls for AI Agents: Stop Your Token Bill From Becoming an Incident
The pattern is now familiar across the industry. A team deploys an AI agent to production —a customer support triage bot, a code review assistant, a research agent that queries internal documentation— and within weeks the monthly invoice arrives showing a five-figure number that nobody budgeted. The agent worked, technically. It returned 200s, finished in seconds, and produced plausible answers. But underneath the smooth operation, runaway tool-call loops, redundant retrievals, and unconstrained reasoning chains quietly burned through the token budget at multiples of the expected rate.
This article explains why traditional observability tooling misses these failures, what changes when you instrument agents specifically, and how teams running agents in production in 2026 are containing both the costs and the silent quality degradation that comes with them. The framing is pragmatic: this is operational practice, not research, and the solutions that work are the ones you can actually deploy.
Why agents fail differently than applications
Conventional observability was built around services that either work or fail. A web server returns 200 or 500. A database query completes in milliseconds or throws an exception. The telemetry stack —Prometheus, OpenTelemetry, the ELK family— was designed to capture these binary outcomes and the latencies that surround them.
AI agents break this model in a specific way. An agent can return HTTP 200, finish in two seconds, and still be wrong. The failure mode is semantic: the agent reached a confident but incorrect answer through a chain of reasoning that looked plausible at every step. As Vellum's 2026 analysis puts it: "an agent can return HTTP 200, finish in two seconds, and be wrong." Traditional monitoring sees a healthy service. The user sees a confidently wrong answer.
This semantic failure mode is what makes agent observability its own discipline. It is not enough to know that the agent was called; you need to know what it did during the call —which tools it invoked, what it retrieved, how it reasoned, where it committed to a decision that turned out to be wrong. That telemetry is structured differently, captured at different points, and analyzed with different questions in mind.
The difference between LLM observability and agent observability
LLM observability, the older discipline, tracks individual model calls. It records the prompt, the completion, the latency, the token count, and the cost. A team running a classification endpoint or a single-turn summarizer can get most of what they need from LLM observability: was the call made, did it succeed, how much did it cost.
Agent observability adds a layer above LLM observability: it tracks the run those calls sit inside. An agent run is a tree of operations: planning calls, routing decisions, tool invocations, retrievals, follow-up LLM calls based on intermediate results, and final synthesis. The interesting failures happen in the middle of this tree, not at the leaves. A tool that returned an empty result is not an error, and is often the reason the final answer was wrong. A retrieval that retrieved the wrong documents is not an exception, and is often the reason the agent hallucinated a citation.
OpenAI's developer community has been collecting these patterns in real time. In a September 2026 thread titled "Has anyone actually solved runaway agent costs?", participants describe the "Zombie Loop" problem —agents that loop on the same tool call indefinitely until they hit a configured maximum, by which time they have consumed thousands of tokens reaching an answer a direct call would have produced. Other participants describe agents that call paid external APIs in tight loops, each call cheap individually but cumulatively devastating. The horror stories all share a pattern: the agent appeared to be working, the telemetry showed healthy service, and the cost only became visible when the invoice arrived.
What a proper agent trace looks like
A trace for an agent run is a tree of spans, each one representing an operation. The root span is the agent run itself; its children are the planning calls, the tool invocations, the retrievals, and the LLM calls. Each span carries its own input, output, latency, and cost. The tree is what tells the story, because an individual span rarely explains itself.
For a research agent that takes a query, plans a retrieval, retrieves documents, plans a synthesis, calls an LLM to draft an answer, and verifies the answer against the retrieved sources, a typical trace might have 10 to 50 spans. Each retrieval has its own query, top-k results, and similarity scores. Each LLM call has its prompt, completion, token count, and finish reason. The synthesis call references the outputs of the prior calls.
When something goes wrong, the trace lets you replay the run step by step. You can see that the retrieval returned documents about the wrong product line, that the agent then synthesized an answer using those wrong documents, and that no verification caught the issue because the verification step used the same flawed retrieval. Without the trace, you would see only "agent returned wrong answer" with no path to remediation.
The OpenTelemetry GenAI semantic conventions
OpenTelemetry has been developing semantic conventions for generative AI workloads under the `gen_ai.*` namespace. These conventions define standardized attribute names for the things that matter in agent instrumentation: the model name, the token counts for prompt and completion, the finish reason, the tool name, the tool arguments, the retrieval query and results.
The conventions are not yet finalized, but they are stable enough that production tooling is adopting them. The benefit of standardization is that instrumentation written once can export to more than one backend. If you instrument your agents using the OpenTelemetry GenAI conventions, you can route the same spans to Datadog, to Honeycomb, to a self-hosted OpenTelemetry collector, or to a specialized agent observability vendor without rewriting your instrumentation.
The auto-instrumentation packages that exist today cover the major frameworks. OpenAI's Python and Node SDKs have community-maintained instrumentations. Anthropic's SDK has official instrumentation. LangChain and LlamaIndex, the two most popular agent frameworks, both ship with OpenTelemetry-compatible instrumentation that captures the major span types automatically. For teams building custom agent frameworks, the OpenTelemetry GenAI conventions provide a template that can be adapted.
The four classes of alerts that actually matter
Standard alerting rules do not work for agents. The traditional approach —alert when latency exceeds a threshold, alert when error rate exceeds a threshold— misses the most common agent failure modes. A latency threshold does not catch an agent that is taking longer because it is looping. An error rate threshold does not catch an agent that is succeeding but returning wrong answers.
The alerting rules that matter in practice are different. OneUptime's 2026 analysis identifies four classes:
Cost rate alerts: "This agent is burning $50/hour, normally it's $3/hour." The threshold is calibrated against historical baseline, not against an absolute number, because agent costs vary wildly by query type. The alert fires when cost-per-time exceeds the baseline by some multiple —two times, five times, ten times depending on tolerance.
Loop detection alerts: "This agent has made 20+ LLM calls for a single query." Most well-functioning agents make fewer than 10 LLM calls per query; alerts on call count catch the zombie loops that burn tokens without making progress.
Latency distribution alerts: Not just p99 latency, but alerts when the variance of latency itself increases. An agent that is taking highly variable time per query is often making inconsistent numbers of internal calls, which often correlates with inconsistent quality.
Quality degradation alerts: "Retrieval relevance scores dropped 40% in the last hour." These alerts require instrumenting the retrieval step specifically and tracking the quality of what was retrieved. They catch the silent quality failures that cost rate alerts miss.
The runaway agent as a budget incident
OneUptime's framing is worth emphasizing: "A runaway agent is a budget incident, and without per-run cost telemetry you find out about it on the invoice." This reframes the problem from engineering to finance. Traditional performance incidents trigger engineering response; budget incidents require finance and procurement involvement. The two are not interchangeable.
For an agent that costs $0.05 per query on average but hits a code path that costs $5 per query due to a loop, the difference between the expected and observed bill can be 100x. If that code path is triggered by a non-trivial fraction of production queries, the monthly bill can exceed the entire engineering budget for the project. Several of the horror stories in the OpenAI developer community involve exactly this pattern: a single non-obvious code path that becomes a recurring cost amplifier.
The mitigation requires per-run cost telemetry. Every agent run needs to carry its total cost as a span attribute, and the cost needs to be computed from the actual token counts and tool invocations, not estimated from call count. This requires instrumenting the cost calculation at every layer of the agent stack, not just at the LLM call level. Tool calls have costs too —API costs for external services, database query costs, retrieval costs from vector databases. An agent that makes 50 retrievals per query can cost more in retrieval fees than in LLM tokens.
Choosing the right tool for the job
The agent observability tool landscape in 2026 splits into several groups. For tracing and evaluation, LangSmith, Langfuse, Arize Phoenix, Braintrust, and Confident AI capture agent runs and score them, with Langfuse and Phoenix favored by teams wanting open-source or OpenTelemetry-native options. For correlating agent behavior with infrastructure, Datadog LLM Observability ties model signals to the rest of the stack. For proxy-based cost and latency tracking, Helicone is quick to deploy. For evaluation and safety, Galileo and Fiddler add guardrails and compliance monitoring for regulated teams.
The choice depends on what you already have. Teams already running Datadog for infrastructure observability will find Datadog LLM Observability the path of least resistance. Teams with OpenTelemetry-native stacks will find Langfuse or Arize Phoenix the natural extension. Teams building evaluation-heavy workflows will find Confident AI or Braintrust the best fit.
A common mistake is to treat the choice as permanent. Agent observability tooling is evolving rapidly, and switching costs are real but not insurmountable if you instrument with the OpenTelemetry GenAI conventions. The conventions are stable enough that switching backends does not require re-instrumenting the agent code.
The cost of not having agent observability
The cost of running an agent without proper observability is not just the runaway token bills. It is the silent quality failures that erode user trust over weeks and months before anyone notices. It is the inability to debug specific failure modes because no trace exists. It is the impossibility of A/B testing new prompts or new agent architectures because the metrics that would distinguish them are not captured.
Teams that ship agents without observability are flying blind. They might get lucky and the agent works well for months. They might also have a quality regression that drives users away, and they will not know why. As OneUptime's analysis concludes: "The AI agent wave is real. The companies that instrument properly from day one will iterate faster, spend less, and ship more reliable products. The ones that don't will be debugging production issues by reading LLM output logs and praying."
The good news is that the tooling has matured enough that instrumenting from day one is no longer a research project. The OpenTelemetry GenAI conventions are stable, the auto-instrumentation packages cover the major frameworks, and the alerting patterns are well understood. The cost of proper instrumentation is days of work, not months. The cost of skipping it is paid in runaway bills, silent failures, and user trust eroded over time.
Conclusión operativa
Instrument your agents with OpenTelemetry GenAI semantic conventions from the start. Capture cost as a span attribute at every layer, not just LLM calls. Set up four classes of alerts: cost rate, loop detection, latency variance, and retrieval quality. Choose an observability backend that fits your stack but stays portable through the conventions. And treat runaway agent costs as budget incidents, not just engineering ones.
The agent wave is not slowing down. The teams that ship reliable agents at predictable cost are the ones that have built the observability discipline to see what their agents are actually doing. Every other team is debugging blind.
---
*Verified sources: InfoQ analysis "Observability and Cost Controls for AI Agents" by Mark Silvester, September 2026; Expanso "AI Agent Observability Best Practices 2026"; OneUptime analysis "Observability for AI Agents: Why Your LLM Apps Are Flying Blind", February 2026; Uptrace guide "OpenTelemetry for AI Systems: LLM and Agent Observability", April 2026; Vellum "Best AI Agent Observability Tools 2026"; Elastic "Observability Trends for 2026 Part 2"; Confident AI "Top 8 AI Agent Observability Platforms for 2026"; OpenAI Developer Community thread "Has anyone actually solved runaway agent costs?", September 2026; Honeycomb "15 Best AI Observability Tools for Production Teams in 2026".*