Three quarters of large enterprises already have AI agents running in production. That is the headline number from Harness’s State of Agent DLC 2026 report, released September 10, and it would sound like a success story — except for the follow-up finding.
When researchers asked those same organisations how confident they were in their ability to test, secure, track, cost, and roll back those agents, confidence scores landed in the mid-70s across every category. That sounds reasonable until you notice that 60% of respondents overran their AI agent budget in the most recent quarter, while 74% claimed to have a complete picture of their per-agent spend. Both statements cannot be true at the same time.
That contradiction is the central finding of a report based on 700 technology professionals at large enterprises across the United States, United Kingdom, France, Germany, and India. Respondents worked in software engineering, IT operations, or technology leadership at organisations with at least 1,000 employees, 100 or more developers, and annual revenue above $100 million. Fieldwork was conducted by Sapio Research in July 2026.
Why Traditional Controls Break Down for Agents
The report’s core argument is that AI agents do not behave like conventional software, and most enterprise governance stacks were built assuming software does.
With deterministic code, the same input produces the same output every time. You can write a test suite, define expected behaviour, and get a pass or fail result with high confidence. AI agents do not work that way. The same agent, given the same starting conditions, can produce meaningfully different outputs across runs. It can call different tools in a different sequence. It can decide to take an action that was never explicitly programmed. The evaluation problem is fundamentally different.
Most organisations are responding to this by applying the same control patterns they already know — version control, integration tests, deployment gates — without adjusting for the non-deterministic nature of what they are governing. The Harness research finds that gap is already showing up in production incidents, security breaches, and the budget overruns mentioned above.
The Five Gaps That Matter
The report identifies five areas where confidence outpaces actual control:
Testing. The majority of organisations say they test their agents before production. Far fewer have evaluation frameworks that account for the range of behaviours an agent can exhibit, the edge cases it may encounter in a live environment, or the ways it can fail silently.
Security. Confidence in agent security runs high, but the questions that matter for agentic systems — what can this agent access, what external services can it call, what does it do if manipulated through injected context — are rarely answered by conventional security reviews designed for APIs.
Inventory. Knowing what agents an organisation is running turns out to be harder than it looks. Business units deploy agents independently. Third-party tools embed agents in their platforms. AI capability spreads faster than governance registers. The report finds significant numbers of organisations that cannot give a complete list of agents currently operating in their environment.
Cost. Token costs, inference infrastructure, and the compounding expense of long-running agent loops create a cost profile that differs sharply from conventional software. The 74% vs 60% contradiction above captures this: organisations believe they have cost visibility because they have billing dashboards, but they lack the per-agent attribution needed to understand which workflows are generating the spend.
Rollback. When a deterministic service misbehaves, rollback is straightforward — redeploy the last known good version. When a long-running agent has been taking actions across multiple systems for hours, the meaning of rollback becomes contested. The research suggests most organisations have not worked through what rollback actually means for their deployed agents.
What This Means for Business Leaders
The Harness report lands at a useful moment. Enterprises are past the pilot phase — the 75% production deployment figure confirms that — but many are discovering that getting agents into production is an easier problem than keeping them under control once they’re there.
There is a risk that organisations respond to early incidents by pulling back on agent deployments rather than improving their governance approach. That would be the wrong lesson. The gap the report identifies is a capability gap, not a technology gap. The agents are working. What is not working is the surrounding infrastructure to monitor, govern, and cost them properly.
The organisations most exposed are those that deployed quickly during the initial wave of enterprise AI enthusiasm, without asking hard questions about how they would know if something went wrong. They now have production agents, some budget overruns, and governance frameworks that were designed for software that behaves differently.
The organisations best positioned are those building control infrastructure alongside deployment — treating agent governance as a first-class engineering problem rather than a compliance exercise applied after the fact.
What This Means for Business
If you are running AI agents in production, or planning to, the Harness findings suggest three practical questions worth asking before the next deployment:
Can you list every agent currently running? Not just the ones IT deployed — the ones marketing built on a no-code platform, the ones customer success started using last month, the agents embedded in the SaaS tools you pay for. If the list is incomplete, the governance picture is incomplete.
What does rollback actually mean in your environment? For agents that take real-world actions — sending emails, updating records, calling external services — rolling back a deployment does not undo what the agent already did. Defining what rollback means, and what acceptable failure looks like, is a prerequisite for operating responsibly.
Are your cost controls agent-native? Per-agent cost attribution is different from per-service cost attribution. If your visibility into AI costs runs through a single billing line, you may have the same confidence-versus-reality gap the report describes.
The report is not an argument against deploying AI agents. It is an argument for being honest about the current state of control infrastructure — and for investing in it with the same seriousness as the deployments themselves.
The legislative response is already arriving. Congress introduced the Stop Rogue AI Act in September 2026, which would require organizations to maintain a machine-readable inventory of every AI agent they run — exactly the kind of visibility the Harness research shows most enterprises currently lack.
At Enterprise DNA, we help businesses build and govern AI agent workforces through Omni Ops and advise on practical AI strategy through Omni Advisory. If you are navigating enterprise AI governance, book a discovery call to talk through your situation.
Source
Harness
Free Resource
Going deeper with Claude?
Get the free 32-page implementation guide for ANZ teams.
Your guide is ready
Check your downloads folder. If it did not open automatically, use the button below.
Download the GuideWant this working inside your business?
See what's possible