Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Latest AI and industry news. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

News Trending AI News

Amazon: Only 5% of AI Agents Actually Reach Production

Amazon's AGI director at VB Transform says 85% of enterprises pilot AI agents but only 5% make production — the gap is reliability, not model smarts.

Enterprise DNA | | via VentureBeat
Amazon: Only 5% of AI Agents Actually Reach Production

At VB Transform 2026 in Menlo Park last week, Bryan Silverthorn, Director of AGI Autonomy at Amazon, delivered one of the more quietly important talks at the year’s leading enterprise AI conference. The message was simple and a little uncomfortable: the reason most companies can’t move their AI agents from pilot to production has nothing to do with model intelligence.

It has everything to do with predictability.

“Enterprises aren’t holding back on AI agents because the technology isn’t smart enough,” Silverthorn said. “They’re holding back because it isn’t predictable enough.”

The data backs this up. Cisco research cited in the presentation shows that 85% of enterprises are actively piloting AI agents. Only 5% have shipped them to production. That is a gap of eighty percentage points, and it has remained stubbornly wide even as the models powering those agents have gotten dramatically better over the past 18 months.

Why the Pilot-to-Production Gap Persists

Silverthorn has an interesting vantage point. He joined Amazon through its acquisition of Adept AI, the startup that was building agents capable of operating desktop software. He now leads multimodal agent training inside Amazon’s AGI lab. He has seen agent failures at scale, which gives his framework credibility.

His core argument is that “reliability” in the context of AI agents is not a single property. It is four distinct things, and treating them as one is part of why enterprise deployments fail.

The four dimensions, which he credited to research from Princeton, are:

Consistency — Does the agent produce the same output for the same input, in the same context, across repeated runs? This sounds basic. It is far from guaranteed with current models.

Robustness — Does the agent degrade gracefully when inputs are messy, ambiguous, or slightly outside the training distribution? Real business data is almost always messy.

Predictability — Can a human operator anticipate what the agent will do next? Can they build intuition about its failure modes? Agents that surprise their operators get shut down. Agents that behave in ways operators understand, even when they occasionally fail, get kept.

Safety — When the agent does fail, how contained is the damage? Can the mistake be caught before it propagates downstream? Can it be reversed?

Agents routinely ace internal evaluations and then, in Silverthorn’s phrase, “collapse in the wild.” The evaluations test capability. The production environment tests reliability. These are not the same test.

Managing Agents Like Interns

The mental model Silverthorn offered for managing agents was deliberately low-tech: call them interns.

Inside Amazon’s AGI lab, researchers apparently do exactly this. They refer to their agents as interns and to multi-agent systems as teams of interns. “I’ll have my intern talk to your intern” is reportedly an actual sentence people say in that building.

The framing is useful because it shifts the operating mindset from software thinking to management thinking. You don’t assume your intern will always do the right thing. You design oversight structures. You ask “what could go wrong here?” You build in checkpoints, escalation paths, and the ability to undo work before it becomes irreversible.

Most enterprise teams deploying AI agents have software engineers in charge of them. Silverthorn’s argument is that what they actually need is managers. The skill of deciding what risk you can accept, defining clear boundaries, and building processes around fallibility is a management skill, not an engineering skill.

The Infrastructure Answer

The practical solution Amazon is building around is architectural: decouple systems and sandbox environments so that even a misbehaving agent cannot cause real damage.

Rather than trusting the model to behave correctly, you design the surrounding infrastructure so that misbehavior has bounded consequences. The agent operates in an environment where its worst-case action is recoverable.

This is a meaningful departure from how most enterprise AI deployments are currently structured. Most deployments give agents access to live systems and trust governance to model-level guardrails. Silverthorn’s point is that model-level guardrails are insufficient when the failure mode is unpredictability rather than outright refusal.

Sandboxing, shadow mode testing, staged rollouts, and comprehensive logging of agent actions are the tools he pointed to. None of these are novel concepts. They are standard practices for deploying any complex software system. They are just underused in agentic AI deployments.

What This Means for Business

The reason this talk matters beyond enterprise architecture circles is what it says about where the real competitive advantage sits in AI deployment right now.

The companies that win with AI agents in 2026 will not necessarily be the ones using the best models. They will be the ones that built the operational infrastructure to run agents in production reliably. The model race is largely irrelevant to that problem. Anthropic, OpenAI, and Google are all producing frontier-grade agents. The differentiator is governance, observability, and workflow design.

For a business owner thinking about deploying AI agents, the VB Transform conversation suggests a reframe. Stop asking “is the AI smart enough to do this?” Start asking “is my operation set up to manage an AI agent doing this?” The answer to the second question almost always reveals more work to do.

The pilot-to-production gap is, at its root, an organisational maturity problem. That is actually good news, because it is solvable without waiting for the next model release.

If you are trying to move AI agents from experiment to a workflow your business actually depends on, book a discovery call with Enterprise DNA. Our Omni team works with businesses to design agent deployments built for production from the start, not retrofitted after the pilot fails.

Working With Claude field guide cover

Free Resource

Going deeper with Claude?

Get the free 32-page implementation guide for ANZ teams.

No spam. Unsubscribe any time.