Enterprise DNA
News Trending Industry

1 in 5 Enterprises Can't Stop Runaway AI Agent Spending

VentureBeat Pulse Research finds 21% of enterprises have no real-time controls on AI agent costs, as Gartner forecasts $207B in agent software spend in 2026.

Enterprise DNA | | via VentureBeat
1 in 5 Enterprises Can't Stop Runaway AI Agent Spending

New research from VentureBeat Pulse has surfaced a governance gap that many enterprise technology leaders will recognise but few have talked about publicly: one in five businesses running AI agents cannot stop a runaway agent’s spending in real time. There is no kill switch. No live cost meter. Just a bill at the end of the month.

The finding sits inside a broader dataset that finds organisations now run three AI orchestration platforms simultaneously on average, yet 21% rely solely on reactive monitoring with no real-time controls in place.

The “Tokenmaxxing” Problem Has a Name Now

The industry has started calling this “tokenmaxxing”: the phenomenon of AI agents consuming massive volumes of tokens without delivering proportionate business value. The pattern usually unfolds the same way. A team deploys an AI coding assistant or a multi-step agent workflow. Usage scales faster than expected. A frontier model handles tasks that a cheaper model could do just as well. No one is watching the meter.

The consequences can be severe. According to the VentureBeat research, Uber gave its engineering teams access to Claude Code in late 2025 and set up internal leaderboards tracking token consumption. By April 2026, the entire AI coding budget for the year was gone. When asked to connect the overblown spending to product improvements for riders and drivers, the company’s President and COO said there was no link yet.

That gap between consumption and output is the heart of the tokenmaxxing problem.

How Enterprises Are (and Are Not) Managing It

The research breaks down the current control landscape:

  • 30% use native budget caps or throttling provided by their AI platform
  • 25% have built custom gateway middleware to enforce policies
  • 25% use dynamic model routing to shift heavy workloads toward lower-cost models
  • 21% have no real-time controls at all

The 21% without real-time controls is the alarming figure, but the other 79% are not in a comfortable position either. Most teams handling model selection are doing it at the prompt level, meaning whoever is prompting makes the call on which model to use. That default behaviour directs substantial spend toward frontier models for tasks where cheaper options would perform equally well.

The Scale of the Problem Is Getting Larger

Gartner expects global spending on AI agent software to reach $207 billion this year, up 139% from $86.4 billion in 2025. If 21% of that enterprise spend cannot be stopped in real time when agents go off-script, the financial exposure across the industry is significant.

This is not a theoretical concern. Token costs have come down considerably through 2026 as providers compete, but the volume of agent activity has grown faster than prices have dropped. For many organisations, the net AI infrastructure bill is rising even as per-token costs fall.

What This Means for Business

If you are deploying AI agents without real-time cost controls, this is a fire hazard, not a hypothetical. Reactive monitoring catches the problem after it has happened. At the rate agents can consume tokens, “after it has happened” can mean hundreds of thousands of dollars in unexpected spend within a billing cycle.

Model routing is underutilised. The 25% of enterprises using dynamic routing to offload tasks to lower-cost models are getting materially better cost efficiency. Most tasks that agents perform, summarisation, classification, basic drafting, data transformation, do not require a frontier model. Routing intelligence can dramatically lower per-task cost without reducing output quality.

The governance layer is not optional. As AI agent deployments scale from pilots to production, the infrastructure around them needs to keep pace. This means policies on which models agents can access, spending thresholds per team or project, audit logging, and the ability to halt a runaway workflow before it becomes a CFO problem.

ROI measurement needs to improve. The Uber example is instructive because the problem was not just overspending, it was overspending without evidence of return. Enterprises that cannot connect AI agent activity to business outcomes are operating blind. Spend without outcome data is impossible to defend at the board level, and increasingly difficult to approve for the next round.

The organisations that are scaling AI agents successfully in 2026 are building governance alongside capability, not as an afterthought. The ones that are not have a cost crisis waiting for them. The difference, according to the VentureBeat data, comes down to whether real-time controls were part of the architecture from the start or bolted on after something went wrong.

Working With Claude field guide cover

Free Resource

Going deeper with Claude?

Get the free 32-page implementation guide for ANZ teams.

Add your name (optional)

No spam. Unsubscribe any time.