Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Latest AI and industry news. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

News AI News

Enterprise AI margins are bifurcating hard: token costs collapsed ~1,000x but agentic consumption is outrunning the savings.

Prices fell from ~$60/M tokens (2021) to ~$0.06/M, yet Anthropic's inference margin climbed from 38% to 70%+ in 2026 on a ~$30B run-rate (80%.

Enterprise DNA |
Enterprise AI margins are bifurcating hard: token costs collapsed ~1,000x but agentic consumption is outrunning the savings.

AI Pulse · Business Models & Winners

The play

Token costs dropped 1,000x but agentic usage is eating the savings, shift budget planning from per-token to per-task or per-agent economics.

Token prices dropped roughly 1,000 times in five years, from about $60 per million tokens in 2021 to six cents today. You’d think that would wreck margins for the companies selling AI services. Instead, the numbers tell a stranger story. Anthropic’s inference margin climbed from 38 percent to over 70 percent in 2026, on a $30 billion run rate that’s 80 percent enterprise. OpenAI sits near 33 percent gross margin on roughly $25 billion annualized revenue.

The reason margins didn’t collapse is that enterprises are using AI differently now. Agentic workflows, where models call other models and loop through tasks, consume tokens at a pace that outstrips the price drop. Telnyx was spending $100,000 a day on Anthropic before switching workloads to Z.AI at $100 per agent per day. Uber burned through its entire 2026 AI budget in four months. The unit cost fell, but the volume exploded.

This creates two problems for operators. First, you can’t budget AI the way you budget SaaS seats. Usage spikes in ways that are hard to predict, especially when you start chaining agents together. Second, the margin split between providers means the market is bifurcating. Some vendors are capturing value by optimizing inference at scale, others are competing on price and losing money on every call.

If you’re rolling out AI internally, you need visibility into what’s actually running and what it costs per task, not per token. That’s the kind of stack management we build into an AI command centre, so you can see where the spend is going before it runs away. The token price drop was real. The consumption explosion is just as real, and it’s the part that will surprise your CFO.

Working With Claude field guide cover

Free Resource

Put what you just read to work

The free 32-page Working With Claude guide: the full ecosystem, Claude Code, and how to roll it out across a business.

No spam. Unsubscribe any time.

Free daily email

Get this every morning.

This brief is one item from today's AI Pulse, the short daily read we run for ourselves on what is actually happening in AI. Subscribe free and it lands in your inbox each morning.

Free daily email

Subscribe to the daily AI Pulse

One short read every morning on what is actually happening in AI. Free.

One email a day. Unsubscribe any time.