Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Thought leadership & research. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

Key Findings

McKinsey finds 60% of AI agent costs go to refinement. Consulting firms should budget 3-5x expected token spend and start with tightly scoped use cases.

Budget 5x Your Token Costs When Piloting AI Agents
Insight ai

Budget 5x Your Token Costs When Piloting AI Agents

Sam McKay

The conversation around AI agents has shifted. Six months ago, every vendor was talking about capability. Now the question is cost.

McKinsey’s latest analysis puts a number on what many consulting firms are discovering the hard way: 60% of AI agent operating costs go to response refinement, not the actual inference work. The token bill you see in your dashboard is a fraction of what you’re really spending when you factor in the engineering time, the prompt iteration, the error handling, and the context management required to make an agent useful in production.

If you’re a partner or GM running a consulting firm and you’ve been piloting AI tools, this matters. The ROI test isn’t passing or failing on whether the agent can draft a proposal. It’s failing on whether you can control the burn rate long enough to prove the value.

Here’s what we’re seeing across firms in our network: the ones that succeed with AI agents start with tightly scoped use cases, budget 3-5x their expected token costs, and treat the first six months as a learning exercise, not a deployment. The ones that struggle try to automate everything at once, underestimate the operational overhead, and pull the plug when the bill surprises them.

This article walks through how to budget for AI agent pilots in a consulting firm, what operating costs actually look like beyond the API calls, and which use cases give you the best chance of proving ROI before you run out of runway.

The Hidden Costs McKinsey Found

The McKinsey report breaks down AI agent costs into three buckets: compute (the token spend), refinement (the engineering and iteration work), and infrastructure (the middleware, monitoring, and orchestration layer). The compute is the smallest bucket.

For a typical consulting firm piloting a proposal generation agent, here’s what that looks like in practice. You pay OpenAI or Anthropic $50 for the tokens to generate a 20-page proposal. That’s the number you see. What you don’t see is the 12 hours your senior associate spent refining the prompt, testing edge cases, and fixing formatting errors. At $150 per hour, that’s $1,800. Then there’s the middleware that pulls in past proposals, manages the context window, and handles the output formatting. That’s another $200 per month in tooling costs, plus the dev time to set it up.

The token cost is 2% of the total. The refinement work is 90%. The infrastructure is the rest.

This pattern holds across every agent use case we’ve built. The API call is cheap. Making the agent reliably useful is expensive. The firms that budget only for tokens get halfway through a pilot, realize they’re burning $10K per month on internal time, and shut it down before they see results.

The ones that budget for the full stack from day one give themselves enough runway to iterate, learn what works, and build something that actually saves time instead of creating a new project.

Start With One Use Case That Hurts

The temptation when you’re piloting AI is to try everything. Proposal generation, research synthesis, meeting summaries, knowledge management. The logic is that if you’re going to invest in the infrastructure, you might as well get leverage across the firm.

That logic is backwards. The infrastructure cost is fixed whether you’re running one agent or five. The refinement cost scales linearly with the number of use cases. Every new agent is another set of prompts to tune, another workflow to map, another set of edge cases to handle.

The firms that control burn rate start with one use case that’s painful enough to justify the investment and narrow enough to prove value in 90 days. For most consulting firms, that’s one of three places: proposal generation, research synthesis, or knowledge management.

Proposal generation is the most visible pain. Every major opportunity requires 20-40 hours of senior time to pull together a tailored pitch. You’re rewriting the same capability descriptions, reformatting case studies, and adjusting pricing for the tenth time this quarter. The work isn’t hard, but it’s expensive and it delays response time.

A Proposal Generation Agent pulls past proposals, case studies, and pricing into a tailored draft for the new opportunity. It doesn’t write the final version, but it cuts the first-draft time from 20 hours to two. The partner still reviews and refines, but the heavy lifting is automated.

Research synthesis is the quieter pain. Every engagement starts with 10-15 hours of secondary research that gets repeated across clients. You’re pulling industry reports, analyzing competitors, and summarizing market trends. The work compounds across the firm because no one’s sharing notes.

A Research Agent runs structured industry and company research at the start of every engagement, with sources, summaries, and a one-page brief. It doesn’t replace the deep expertise your team brings, but it eliminates the repetitive legwork and gives everyone a common starting point.

Knowledge management is the long-term pain. Every project produces IP. Almost none of it is reusable across the firm. You’ve got hundreds of decks, docs, and meeting transcripts sitting in SharePoint, and the only way to find anything is to ask someone who was there.

A Knowledge Agent reads every deck, doc, and meeting transcript the firm produces and answers questions across the corpus. It’s not a search tool, it’s a synthesis tool. You ask it how you’ve approached pricing in past healthcare engagements, and it pulls examples, summarizes patterns, and gives you a starting point.

Pick one. Budget for it properly. Prove it works. Then add the next one.

What 3-5x Your Token Budget Actually Means

Here’s a worked example. You’re piloting a proposal generation agent. You estimate you’ll generate 20 proposals per month. Each proposal costs $50 in tokens. Your monthly token budget is $1,000.

Now add the refinement work. You’ll spend 40 hours in month one tuning the prompts, mapping the workflow, and testing edge cases. That’s $6,000 in internal time. Months two and three drop to 10 hours each as you stabilize the system. That’s another $3,000.

Add the infrastructure. You need middleware to pull in past proposals, manage the context window, and format the output. That’s $500 per month in tooling costs, plus 20 hours of dev time to set it up. That’s another $3,500 in month one, $500 per month after.

Total cost for the first three months: $1,000 in tokens, $9,000 in refinement, $4,500 in infrastructure. Your token budget was $3,000 for the quarter. Your actual cost was $14,500. That’s 4.8x.

This isn’t a worst-case scenario. This is typical for firms that take the pilot seriously and invest the time to make the agent useful. The firms that undershoot the budget either cut corners on refinement (and end up with an agent no one trusts) or burn out their team trying to make it work on nights and weekends.

The 3-5x multiplier gives you room to iterate, fix mistakes, and build something that actually works. It’s not waste, it’s the cost of learning.

The Omni Audit Finds Your Highest-ROI Agent

We built the AI audit for consulting firms because most firms don’t know where to start. You’ve got a dozen painful processes, limited budget, and no clear sense of which agent will pay back fastest.

The Omni Audit is 60 minutes. We walk through your proposal workflow, your research process, and your knowledge management setup. We map the time cost, the error rate, and the downstream impact of each pain point. Then we give you three outputs: a ranked list of agent opportunities, a 90-day pilot plan for the top use case, and a budget breakdown that includes tokens, refinement, and infrastructure.

No deck. No follow-up meeting. Just a clear recommendation you can act on.

The firms that book a 60-min Omni Audit come in with a vague sense that AI should help and leave with a specific plan, a realistic budget, and a use case that’s worth the investment. Most of them start the pilot within two weeks.

If you want a structured way to think through your first agent deployment before the audit, we’ve built a worksheet that walks through the same prioritization framework we use. Deploy Your First Business Agent is a one-page checklist that helps you map your highest-cost processes, estimate the automation potential, and calculate the payback period. It’s the same tool we use internally when we’re scoping a new agent build.

Why Consulting Firms Leak $80K-$300K Per Year on Repeated Work

The McKinsey finding about operating costs is part of a bigger story. AI agents are expensive to run, but the work they’re replacing is more expensive.

A mid-sized consulting firm doing $5M in revenue typically has three to five partners and 10-15 associates. Each partner spends 15-20% of their time on proposals, research, and internal knowledge management. That’s 300-400 hours per year, per partner. At $300 per hour, that’s $90K-$120K per partner in opportunity cost.

Multiply that across the firm and you’re looking at $270K-$600K in senior time spent on work that could be systematically automated. Not all of it, but enough that a well-designed agent could recover 30-50% of those hours.

The firms that treat AI pilots as a cost center miss this. They see the $15K pilot budget and compare it to the $3K they were expecting. They don’t compare it to the $150K they’re spending every year on repeated work.

The ROI case for AI agents in consulting isn’t about eliminating headcount. It’s about giving your senior people their time back so they can do the work that actually differentiates the firm. The proposal generation agent doesn’t replace the partner, it gives the partner 60 hours per quarter to spend on client work instead of formatting decks.

That’s the shift McKinsey is describing. The operating costs are real, but they’re a rounding error compared to the cost of not automating.

What We’re Building at Omni

We built Omni because the gap between AI capability and AI usefulness is almost entirely an operations problem. The models are good enough. The tooling is good enough. What’s missing is the middleware layer that connects the model to your actual workflow, handles the edge cases, and makes the agent reliable enough to trust.

For consulting firms, that means three things. First, agents that pull context from your existing systems without requiring you to rebuild your tech stack. The Proposal Generation Agent reads your past proposals directly from SharePoint or Google Drive, no migration required.

Second, agents that handle the refinement work internally so you’re not burning senior time on prompt tuning. The Research Agent comes with pre-built workflows for industry research, competitor analysis, and market synthesis. You configure it once, it runs the same way every time.

Third, agents that get smarter as you use them. The Knowledge Agent learns from every document you feed it and improves its answers over time. You’re not maintaining a static knowledge base, you’re building a system that compounds.

We’re not trying to sell you a platform. We’re trying to help you deploy one agent that works well enough to prove the ROI case, then build the next one. The firms that succeed with AI don’t buy a suite of tools and hope for the best. They start with one painful process, automate it properly, and expand from there.

If you want to see what that looks like for your firm, book my Omni Audit. We’ll map your highest-cost processes, identify the agent with the best payback, and give you a realistic budget that accounts for the full operating cost. No deck, no sales pitch, just a clear plan you can act on.

The firms that control AI agent costs don’t spend less. They spend smarter. They budget for the full stack, start with a tightly scoped use case, and give themselves enough runway to learn what works. That’s the difference between a failed pilot and a system that saves 200 hours per quarter.

The token bill is the easy part. The refinement work is where you prove whether AI is worth the investment. Budget for it properly, and you’ll have room to build something that actually works. Underestimate it, and you’ll shut the pilot down before you see results.

For more on how consulting firms are approaching AI strategy, explore our insights library or dive into the broader Omni platform we’ve built to support operational AI deployments. If you’re earlier in the learning curve, our guides section covers the foundational concepts that make agent deployments successful.