Key Takeaways
- Uber exhausted its entire 2026 AI budget by April, four months in, after Claude Code adoption outpaced its finance models (Forbes, corroborated across multiple outlets).
- Token consumption can vary up to 30x on identical tasks, and models struggle to predict their own eventual cost (Stanford Digital Economy Lab, April 2026).
- Dashboards and spend visibility are now common. Enforcement, stopping the next expensive call before it fires, is the part most teams haven't built (The Pragmatic Engineer, April 2026).
- Closing the gap requires three things to hold at once, before the call goes out: which agent is acting, what budget remains under its identity, and whether the action is permitted.
Uber exhausted its entire 2026 AI budget by April, four months into the calendar year, after Claude Code spread across roughly 5,000 engineers faster than the company's finance models had anticipated. At that speed, scale, and volume, spend has to be watched on purpose, continuously. Watching after the fact is what got Uber to April.

Adoption moved faster than the finance models built to track it. Source: Forbes, corroborated by Storyboard18.
Agentic AI has made cost a runtime variable
The cost of an agent isn't known when the request is made. Agentic tasks can consume wildly different amounts of resources for the same objective, and the gap isn't small. A Stanford Digital Economy Lab study analyzing trajectories from eight frontier LLMs on SWE-bench Verified found that runs on the same task can differ by up to 30x in total tokens. The same study found that models systematically struggle to predict their own eventual consumption.
Part of the reason is structural, not just noisy. Agentic workflows can recursively call models and tools, expanding their own execution path as they go. A task that looks like one request can branch into several without anyone deciding it should.
One user action, many agents, one number at the end
That same Stanford research found that task difficulty, as rated by human experts, only weakly aligns with actual token costs. There's a real gap between how complex a task looks to a person and how much computational effort an agent actually spends solving it.

In practice, one user instruction can trigger planning calls, retrieval, browser actions, code execution, retries, model escalation, and delegation to other agents. The final bill may be perfectly accurate at the account level. It's also close to useless for answering which branch of that execution actually caused it.
Attribution is the other half of control
Follow one request through a realistic chain. A customer request initiates Agent A. Agent A invokes Agent B. Agent B calls a retrieval service and a frontier model. A retry triggers another model call. The invoice sees tokens. The business sees one customer request. Neither one tells you which agent, or which branch of that chain, actually created the cost.
A ceiling can stop the number from growing. It can still leave the harder question sitting there unanswered: which agent, which owner, which feature actually caused it. Prevention without attribution is control with no accountability behind it, a circuit breaker with no name attached to what tripped it. (This is the same identity gap covered in more depth in Ungoverned Agents: The Authorization Gap: an authentication record isn't automatically an actor record, for spend any more than for access.)
The industry has largely solved visibility. It hasn't solved control.
Spend dashboards, cost breakdowns by model and team, numbers that update in something close to real time: this exists widely now. The Pragmatic Engineer documented teams building exactly this kind of tooling in April 2026, spot-checking heavy users and adjusting model defaults. One infrastructure company described its own approach plainly: "We're monitoring but not restricting."
That's the tell. Monitoring is not the same as preventing the next expensive request. The assumption that someone is watching closely breaks down the moment many things start spending independently, and by the time a human notices abnormal spend, the agent may have already finished the expensive execution that caused it.
The distinction that actually matters: a budget alert reports that spend crossed a threshold, after the fact. An enforcement mechanism prevents the next operation from proceeding, before it happens. Those are different categories of tool, not two points on the same maturity curve.
What actually closes the gap
Cost gets checked before the call goes out. A number known only after the tokens are spent is a transcript, not a control.
Identity has to travel with the call chain itself, from Agent A to Agent B to whatever model it invokes, carrying the originating request through every hop, separate from whichever credential happens to be footing the bill. The decision has to sit inside the execution path, not beside it: a gate evaluates the next call before it fires, while a dashboard only ever describes the last one.
All three of these have to hold at once, before the call goes out: which agent is acting, what budget remains under its identity, and whether this specific action is permitted. Miss any one of them, and what's left is observation with better instrumentation, not control.
This is the model Forgebench is built around: a credential that carries identity through the full call chain, a ceiling checked before the call fires rather than a report after it, and a record that ties every dollar back to the agent, owner, and feature that spent it. Spend visibility was never the hard part. Knowing who to ask before the money is gone is.


