Why Your AI Agent Budget Is 3x What You Think It Is
McKinsey just published research that should change how every consulting firm budgets for AI agents. Sixty percent of agentic AI costs go to response refinement. Not the initial query. Not the data retrieval. The back-and-forth cycles where the agent checks its work, validates sources, and iterates toward something you’d actually send to a client.
If you’re planning to automate proposal generation or research synthesis, multiply your expected API costs by three. That’s the real number. And for infrequent tasks, manual review might still be cheaper than running an agent through five refinement loops.
This matters because most consulting firms are entering 2026 with a clear mandate to deploy AI agents for client deliverables. The promise is simple: cut 20 hours off every major proposal, stop re-researching the same industries, turn your knowledge base into something that actually answers questions. But the economics are messier than the vendor pitch decks suggest.
I’m Sam McKay, founder of Enterprise DNA. We’ve spent the last 18 months building Omni, an AI operating system for professional services firms. We’ve deployed agents that write proposals, run research briefs, and synthesize knowledge across hundreds of past engagements. And we’ve learned the hard way that the cost structure of agentic AI doesn’t match what most firms expect when they sign the first contract.
Here’s what the real economics look like, why refinement costs dominate, and how to decide whether an agent is cheaper than the person doing the work today.
The Hidden Cost Layer in Every AI Agent
When you ask an AI agent to draft a proposal, you’re not paying for one API call. You’re paying for a workflow that looks more like this:
- Initial retrieval: Pull relevant past proposals, case studies, and pricing models from your knowledge base.
- First draft: Generate a proposal structure and populate sections.
- Validation loop: Check that the case studies match the prospect’s industry, verify pricing is current, confirm the scope aligns with what your team actually delivers.
- Refinement one: Adjust tone to match your firm’s voice, tighten the executive summary, add specific metrics from past client work.
- Refinement two: Cross-check that no outdated service lines are referenced, ensure compliance language is current, validate that named team members are still with the firm.
- Final pass: Format for your template, add visual elements, generate a one-page summary.
Each of those steps is a separate set of API calls. The validation and refinement loops are where costs pile up, because the agent is reading more tokens than it writes. It’s checking context, comparing outputs, and running conditional logic. McKinsey’s 60% figure reflects this reality: most of the compute budget goes to making sure the output is good enough to use.
For a consulting firm, this means the $200 you budgeted for API costs on a proposal agent is actually $600. And if the agent needs three tries to get the tone right, you’re at $800. That’s still cheaper than 20 hours of senior consultant time, but it’s not the $50-per-proposal number you saw in the demo.
When Manual Review Is Cheaper Than Agent Refinement
Here’s the uncomfortable truth: for tasks you run less than twice a month, paying a human to review and edit a rough draft is often cheaper than running an agent through enough refinement loops to produce client-ready work.
Let’s use research synthesis as an example. Your firm kicks off a new engagement and needs a 10-page brief on the client’s industry, competitive landscape, and regulatory environment. A research agent can pull sources, summarize findings, and structure the brief in about 30 minutes. But if the brief needs to be polished enough to include in your onboarding deck, the agent will run validation loops to check source credibility, cross-reference claims, and adjust the narrative flow.
If you’re only running this task twice a month, the refinement cost per brief might hit $400. A junior consultant can review and polish the rough draft in two hours for $300 in fully loaded cost. The agent is faster, but it’s not cheaper until you’re running the task weekly.
This is why the AI audit for consulting firms starts with frequency mapping. We list every task you’re considering for automation, estimate how often it runs, and calculate the breakeven point. For high-frequency work like proposal generation or client status updates, agents win immediately. For infrequent work like annual strategy decks, manual review still makes sense.
The McKinsey research reinforces this. If 60% of your agent cost is refinement, you need enough volume to amortize that overhead. One-off tasks don’t hit that threshold.
What Refinement Costs Look Like in Practice
We built a Proposal Generation Agent for a strategy consulting firm that writes 40 proposals per quarter. The firm expected API costs around $150 per proposal based on token counts for the initial draft. Actual costs settled at $520 per proposal after three months of production use.
The breakdown:
- Initial draft: $120 in API calls to retrieve past work and generate structure.
- Validation loop: $180 to check case study relevance, verify pricing, and confirm scope accuracy.
- Tone refinement: $140 to adjust voice, tighten language, and match the firm’s editorial standards.
- Final formatting: $80 to apply templates and generate the executive summary.
The firm is still saving money. A senior consultant was spending 22 hours per proposal at a fully loaded cost of $4,800. The agent produces a draft in 90 minutes that needs four hours of partner review. Total cost per proposal is now $1,720 including API spend. That’s a $3,000 saving per proposal, or $120,000 per quarter.
But the economics only work because the firm writes 40 proposals per quarter. If they wrote five, the agent setup cost and ongoing refinement budget wouldn’t pay back.
The Three Tasks Where Refinement Costs Are Worth It
Not every consulting task benefits from AI agents, even when the technology works. Based on what we’ve deployed across firms in the $1M to $25M range, three tasks consistently justify the refinement overhead:
Proposal generation. If your firm writes more than 10 proposals per quarter, a Proposal Generation Agent pays back in the first 60 days. The agent pulls past proposals, case studies, and pricing models into a tailored draft. Refinement loops validate that the content matches the opportunity and adjust tone to fit your firm’s voice. We typically see 18 hours saved per proposal after accounting for partner review time.
Research synthesis. At the start of every engagement, your team runs secondary research on the client’s industry, competitors, and market position. A Research Agent structures this work into a one-page brief with sources and summaries. Refinement loops check source credibility and cross-reference claims. For firms running more than eight engagements per quarter, the agent cost is lower than junior consultant time.
Knowledge base queries. Every project your firm completes generates IP: frameworks, models, slide decks, meeting notes. Almost none of it is reusable because no one can find it. A Knowledge Agent reads your entire corpus and answers questions across past work. Refinement loops validate that the answer is current and relevant to the query. This one is harder to quantify, but firms report saving 6 to 10 hours per week in “someone must have done this before” searches.
These three tasks share a pattern: high frequency, structured inputs, and clear quality standards. That combination makes refinement costs predictable and justifiable.
If you want a practical framework for deciding which tasks to automate first, we’ve published a worksheet that walks through frequency mapping, cost estimation, and breakeven analysis. You can grab it here: Deploy Your First Business Agent. It’s the same tool we use in Omni Audits to prioritize agent deployment.
How to Budget for Refinement Without Blowing Up Your P&L
The McKinsey finding that 60% of costs go to refinement gives you a simple budgeting rule: take your expected API cost and multiply by 2.5 to 3. That’s your real cost per task.
Here’s how to apply that in practice:
Start with token counts. Most AI vendors will give you token estimates for a sample task. If the initial draft costs $100 in API calls, budget $250 to $300 for the full workflow including refinement.
Track actual costs for 30 days. Your first month in production will show you where refinement loops are expensive. Some tasks need more validation than others. Adjust your budget based on real data, not vendor estimates.
Set a cost ceiling per task. Decide the maximum you’ll spend on agent refinement before you revert to manual work. For proposals, we typically see firms set a ceiling at $600 per draft. If refinement pushes past that, the task goes back to a human.
Run a quarterly cost review. Agent costs drift as your team asks for more refinement, adds new validation steps, or expands the knowledge base. Review actual spend every 90 days and adjust your task list. Some tasks will drop off because they’re not hitting volume. Others will scale up because the economics improved.
The firms that manage agent costs well treat refinement as a variable cost tied to task frequency. They don’t budget a flat monthly fee. They budget per task and track whether the economics still work.
What an Omni Audit Tells You About Your Agent Economics
When we run an Omni Audit for a consulting firm, we’re answering three questions:
- Which tasks are you running frequently enough to justify agent refinement costs?
- What does your current manual cost look like for those tasks, including hidden time?
- Where’s the breakeven point for each agent, and how long until payback?
The audit takes 60 minutes. No deck, no follow-up meeting. You walk away with a task priority list, a cost model for each agent, and a 90-day deployment plan.
We’ve run this audit for 40+ consulting firms in the last year. The pattern is consistent: firms overestimate how often they run low-frequency tasks and underestimate the refinement cost for high-frequency work. The audit fixes both problems.
If you’re planning agent deployment in the next quarter, book a 60-min Omni Audit and we’ll map your task list to real economics. You’ll know exactly where agents pay back and where manual work is still cheaper.
The Refinement Tax Is Real, But It’s Not a Dealbreaker
McKinsey’s research on agentic AI costs should make you more careful about budgeting, not more skeptical about deployment. Sixty percent of costs going to refinement is a feature, not a bug. It’s what makes the output good enough to send to a client.
The firms that succeed with AI agents in 2026 will be the ones that budget for refinement upfront, track costs by task, and adjust their deployment plan based on real economics. The firms that struggle will be the ones that expect API costs to match the demo and don’t account for the validation loops that make agents useful.
For consulting firms in the $1M to $25M range, the annual leakage from manual proposal work, repeated research, and inaccessible knowledge sits between $80,000 and $300,000. That’s the cost of doing this work the old way. Even with refinement overhead, agents capture most of that leakage if you deploy them on the right tasks.
We’ve built Omni to handle the full agent lifecycle: task mapping, cost modeling, deployment, and ongoing refinement tuning. If you want to see what that looks like for your firm, book my Omni Audit and we’ll walk through your task list together.
The economics are messier than the pitch decks suggest. But the payback is real if you budget for the refinement tax and deploy agents where frequency justifies the cost. That’s the lesson from McKinsey’s research, and it’s what we’re seeing across every consulting firm we work with.
For more on how AI agents fit into the broader automation strategy for professional services, check out the insights library or explore the Omni Ops platform that powers the agents we’ve discussed here.