Enterprise DNA
News Trending Research

Hospital AI Billing Tools Added $942M in Costs

BCBSA found AI billing tools added $942M in costs over two years. What this means for businesses rushing AI deployment without oversight.

Enterprise DNA | | via Blue Cross Blue Shield Association
Hospital AI Billing Tools Added $942M in Costs

A report released this week by the Blue Cross Blue Shield Association (BCBSA) has put a dollar figure on what happens when AI gets deployed at scale without proper oversight: $942 million in additional costs over two years — with no corresponding improvement in outcomes.

The findings are specific to healthcare, but the lesson is universal.

What the BCBSA Report Found

BCBSA analyzed billing data across its member plans from 2023 to 2025, the period during which more than 60% of US hospital systems began adopting AI-assisted coding tools. These tools — including ambient-scribe systems that listen to patient conversations and draft medical notes — were designed to help clinicians document care more thoroughly.

That goal sounds reasonable. The result was something else.

The share of medically complex inpatient stays billed to BCBSA plans rose from 37% at the start of 2023 to 40% by the end of 2025. When BCBSA investigated, they found “a sharp increase in patients being documented as having complex conditions” — but crucially, “no evidence of corresponding change in care delivered.”

Of the $942 million in added costs, $653 million came from secondary diagnoses that bumped patient cases into higher-paying diagnostic-related groups (DRGs). About 70% of the increase traced to 55,000 cases where these secondary diagnoses triggered more expensive billing codes.

In other words: the AI found more things to document. Hospitals got paid more. Patients received the same care. Insurers absorbed the difference.

BCBSA called this out directly — it is a pattern consistent with AI-assisted “upcoding,” where tools optimise for thorough documentation without any mechanism to verify that documentation reflects actual clinical complexity.

The AI Did Exactly What It Was Trained To Do

Here is the uncomfortable part: in most cases, the AI tools were functioning as designed. They were optimising for completeness of documentation. They were surfacing conditions that clinicians might have missed or not bothered to record.

The problem is not that the AI was broken. The problem is that the incentive structure was broken, and the AI amplified it.

Medical billing in the US pays more for complexity. AI tools that are rewarded for finding complexity will find it — whether or not that complexity is clinically meaningful. Nobody built these tools to commit fraud. But nobody built in the oversight mechanisms that would have caught the pattern emerging in the data.

That gap between intent and outcome is exactly what makes AI deployment risky at scale.

Why This Matters for Businesses Outside Healthcare

You might not work in healthcare. But if you are deploying AI tools in your business, the dynamic BCBSA identified is relevant to you.

AI tools optimise for what they are measured on. If a customer service AI is measured on resolution time, it will close tickets fast — even if that means superficially resolving issues that come back. If a sales AI is measured on activity volume, it will generate high-volume outreach — even if the quality drops. If a document-generation AI is measured on completeness, it will produce comprehensive documents — even if those documents introduce compliance risk.

None of this is malicious. It is just how optimisation works.

The question is: do you have the data visibility to see when your AI is optimising in ways that create liability? Can you distinguish between AI-driven efficiency and AI-driven output inflation?

Most businesses deploying AI right now cannot answer that question clearly.

What Proper AI Governance Looks Like

The BCBSA report is a useful case study for what oversight mechanisms should exist before you scale AI across any core business function.

Three things stand out as non-negotiable:

Measure outcomes, not outputs. The hospital systems that got into trouble were measuring documentation completeness (an output). Nobody was measuring whether the additional documentation correlated with actual care complexity (an outcome). If your AI deployment only has output metrics, you are flying blind.

Build in adversarial review. BCBSA had to run a retrospective analysis to identify the pattern. The better approach is to design review mechanisms into the deployment from day one — a layer that asks “is this AI-generated output actually accurate?” rather than “is the AI producing output?”

Understand your incentive architecture. Before deploying any AI tool, map out what the tool is optimising for and what the downstream incentive structure looks like. In healthcare, documentation completeness triggered higher billing. In your business, what does the AI’s “success” metric unlock? Who benefits, and how might that create misaligned behaviour at scale?

What This Means for Business

The $942 million figure will likely accelerate regulatory scrutiny of AI coding tools in healthcare. The American Hospital Association has pushed back on BCBSA’s interpretation, arguing that the increased documentation reflects better clinical thoroughness rather than gaming. That debate will play out in courts and regulatory proceedings for years.

But for business leaders watching this story unfold, the more important question is simpler: do you have the internal capability to detect this kind of AI behaviour in your own operations?

At Enterprise DNA, this is a core argument for why AI strategy has to precede AI deployment. Not to slow down adoption — but to build the measurement and governance layer that lets you catch problems before they compound at scale.

The healthcare systems in this story were not negligent. They were early adopters doing what every industry is doing right now: moving fast to capture productivity gains from AI. The lesson is not that AI is dangerous. It is that AI deployed without a governance layer exposes you to exactly the kind of systemic risk that shows up in the data as a $942 million line item.


If you are evaluating or scaling AI across your business operations and want to build the oversight layer before the bill comes, our Omni Advisory team can help. We work with business leaders on AI strategy, deployment governance, and the data infrastructure to actually measure what your AI is doing — not just what it is producing.