AI Agent Operating Costs Will Eat Your Consulting Margin
McKinsey just published a warning that most consulting firms deploying AI agents haven’t heard yet: the operating costs of running agents at scale can destroy your ROI faster than the time savings create it. Inference costs, refinement loops, and the hidden expense of keeping agents accurate compound quickly once you move past proof-of-concept and into production delivery.
If you’re a partner or owner at a consulting firm thinking about agents for proposal generation, research synthesis, or knowledge management, you need to model the per-engagement operating cost before you commit budget or promise clients agent-driven deliverables. The math changes fast when you scale from ten proposals a quarter to fifty, or when your research agent runs deep industry scans for every new engagement.
I’m Sam McKay, founder of Enterprise DNA. We’ve built and deployed business agents for consulting firms doing $1M to $25M in revenue. The firms that get ROI right model operating costs up front and tie agent deployment to specific, repeatable workflows where the time saved is measurable and the inference expense is predictable. The firms that don’t often discover six months in that their agent is burning $4,000 a month in API calls and still producing outputs that need two hours of partner review.
This article walks through the operating cost reality for three high-value agent use cases in consulting firms, shows you what sustainable deployment looks like, and gives you a framework to model ROI before you build or buy. If you’re evaluating agents for your firm, this is the conversation you need to have with your team before you sign a contract or allocate engineering time.
The McKinsey Warning: Inference Costs Compound Faster Than You Think
The recent McKinsey report on AI agent economics makes one point very clear: enterprises are underestimating the operating costs of running agents in production. Inference costs (the expense of each API call to the underlying model), refinement loops (when the agent needs multiple passes to get output right), and the human review layer all add up quickly once you move from pilot to scale.
For consulting firms, this matters because the three workflows where agents deliver the most value are also the three workflows where operating costs can spiral if you don’t design the system correctly. Proposal generation, research synthesis, and knowledge management all involve large context windows, multiple model calls per task, and outputs that need to be accurate enough to put in front of clients or use in billable work.
A proposal generation agent that pulls from your past proposals, case studies, and pricing models might make fifteen to twenty API calls to assemble a single draft. If each call costs $0.02 to $0.08 depending on the model and token count, and you’re generating fifty proposals a quarter, you’re looking at $150 to $400 in direct inference costs per quarter just for that one workflow. Add refinement (when the agent needs to rewrite sections based on feedback), and the number doubles.
That’s not catastrophic, but it’s also not free. And if you haven’t modeled it, you don’t know whether the time saved justifies the cost. A senior consultant who used to spend twenty hours on a major proposal and now spends four hours reviewing an agent draft has saved sixteen billable hours. At $200 per hour, that’s $3,200 in recovered capacity. If the agent cost $80 to run, the ROI is obvious. If it cost $800 because the system wasn’t optimized, the math gets harder.
The firms we work with that get this right treat operating costs as a design constraint from day one. They pick workflows where the time saved is large and repeatable, they optimize the agent to minimize unnecessary API calls, and they measure both the cost and the time saved per task. The firms that don’t often discover the problem when the invoice arrives.
Three High-Value Workflows Where Operating Costs Matter Most
Let’s walk through the three workflows where consulting firms see the most value from agents, and where operating costs need to be modeled carefully.
Proposal Generation: Saving Partner Time Without Burning Budget on Rewrites
Every major proposal in a consulting firm starts the same way. A partner or senior consultant opens a blank deck, pulls up three past proposals that might be relevant, copies a few slides, rewrites the executive summary, updates the pricing, and spends twenty to forty hours assembling a document that’s 60% recycled and 40% new.
The time cost is brutal. If your firm writes ten major proposals a quarter and each one takes thirty hours of senior time, that’s 300 hours per quarter that could be spent on billable client work or business development. At $200 per hour, that’s $60,000 in opportunity cost every quarter.
A proposal generation agent built on Omni Ops pulls past proposals, case studies, client testimonials, and pricing models from your firm’s knowledge base and assembles a tailored draft in minutes. The partner reviews the draft, edits the sections that need customization, and submits a polished proposal in four to six hours instead of thirty.
The time saved is real. But the operating cost depends on how the agent is designed. If the agent loads your entire proposal library into context every time it runs, you’re paying for thousands of tokens per call. If it uses a retrieval system that pulls only the relevant past proposals based on the new opportunity, the token count drops by 80% and so does the cost.
We’ve seen firms reduce proposal prep time by 70% with a well-designed agent. We’ve also seen firms spend $1,200 a month on inference costs because the agent was making redundant calls or loading too much context. The difference is in the design, and the design starts with modeling the operating cost per proposal before you build.
If you’re evaluating agents for proposal generation, the AI audit for consulting firms walks through the cost model, the workflow integration, and the expected ROI in sixty minutes. You’ll leave with a cost estimate, a deployment plan, and a clear answer on whether the math works for your firm.
Research Synthesis: Structuring Secondary Research Without Repeating Work
Every consulting engagement starts with research. Industry trends, competitor analysis, regulatory context, financial benchmarks. The work is necessary, but it’s also repetitive. A consultant doing research for a healthcare client in Q1 and a different healthcare client in Q3 is often pulling the same sources, reading the same reports, and synthesizing the same trends.
The time cost compounds across the firm. If each engagement requires twenty hours of secondary research and your firm runs forty engagements a year, that’s 800 hours of research time annually. At $150 per hour, that’s $120,000 in cost that’s largely duplicated across projects.
A research agent built on Omni Ops runs structured research at the start of every engagement. It pulls industry reports, competitor filings, news archives, and regulatory updates, synthesizes the findings into a one-page brief with sources, and delivers it to the engagement team in hours instead of days. The consultant reviews the brief, adds client-specific context, and moves directly into analysis.
The operating cost for a research agent depends on how much data it needs to process per engagement and how many refinement loops it runs. A research task that involves scanning fifty documents and summarizing ten key trends might cost $5 to $15 in inference fees. If the agent saves fifteen hours of consultant time at $150 per hour, that’s $2,250 in recovered capacity for $15 in operating cost. The ROI is clear.
But if the agent isn’t designed to filter sources efficiently, or if it runs multiple passes because the prompt isn’t specific enough, the cost can double or triple. We’ve worked with firms where the research agent was making redundant API calls because the retrieval system wasn’t optimized. The fix took two days and cut operating costs by 60%.
The key is to model the cost per engagement before you deploy. If you’re running forty engagements a year and the agent costs $10 per engagement, that’s $400 annually in operating costs. If it saves 600 hours of consultant time, the ROI is 375 to 1. But you need to know the numbers before you commit.
Knowledge Management: Making Past Work Reusable Without Manual Tagging
Consulting firms produce enormous amounts of intellectual property. Every engagement generates decks, reports, meeting notes, research summaries, and client deliverables. Almost none of it is reusable across the firm because it’s stored in folders, never tagged, and impossible to search effectively.
The cost is invisible but massive. When a consultant needs to find a past analysis of supply chain risk in manufacturing, they either spend two hours digging through shared drives or they redo the analysis from scratch. The firm has already paid for that insight once. Now it’s paying again.
A knowledge agent built on Omni Ops reads every document the firm produces, indexes it, and answers questions across the entire corpus. A consultant can ask “What did we recommend for supply chain risk in the automotive sector?” and get a summary with links to the relevant past work in seconds.
The operating cost for a knowledge agent is tied to the size of the corpus and the frequency of queries. Indexing 10,000 documents might cost $200 to $500 in one-time inference fees. Answering queries costs $0.05 to $0.20 per query depending on the complexity. If your firm runs 500 queries a month, that’s $25 to $100 in monthly operating costs.
The ROI depends on how much time the agent saves. If each query saves thirty minutes of search time, and you’re running 500 queries a month, that’s 250 hours saved per month. At $150 per hour, that’s $37,500 in recovered capacity for $100 in operating costs. The math works.
But if the agent isn’t designed to retrieve only the relevant documents, or if it’s re-indexing the corpus every time someone asks a question, the operating cost can spiral. We’ve seen firms where the knowledge agent was costing $800 a month because the indexing process wasn’t optimized. The fix was straightforward, but it required modeling the cost up front.
If you want to see how a knowledge agent would work for your firm, Book a 60-min Omni Audit. We’ll map your document corpus, estimate the indexing cost, and show you what the query interface looks like in practice.
How to Model Operating Costs Before You Deploy
The firms that get agent ROI right model operating costs before they deploy, not after. Here’s the framework we use with consulting firms evaluating agents for the first time.
Start with the workflow, not the technology. Pick one repeatable task where the time cost is measurable and the output is consistent. Proposal generation, research synthesis, and knowledge management are all good candidates because the work is repetitive and the time saved is easy to quantify.
Estimate the token count per task. Every API call to an AI model costs money based on the number of tokens processed. A token is roughly four characters. If your proposal generation agent needs to read five past proposals (10,000 words total) and generate a new proposal (3,000 words), that’s roughly 16,000 tokens in and 4,000 tokens out. At current pricing for most enterprise models, that’s $0.30 to $1.20 per proposal depending on the model.
Model the frequency. If you’re generating fifty proposals a quarter, and each one costs $0.80 in inference fees, that’s $40 per quarter in operating costs. If each proposal used to take thirty hours of senior time and now takes four hours, you’ve saved twenty-six hours per proposal, or 1,300 hours per quarter. At $200 per hour, that’s $260,000 in recovered capacity for $40 in operating costs.
Add refinement loops. Most agents don’t get the output right on the first pass. If your proposal generation agent needs two refinement passes per proposal (where the user gives feedback and the agent rewrites sections), the token count doubles and so does the cost. That’s still $80 per quarter for $260,000 in time saved, but it’s a number you need to know.
Factor in human review time. An agent that saves you twenty-six hours per proposal but produces output that needs six hours of review is still saving twenty hours. But if the output needs twelve hours of review because the agent isn’t accurate, the time saved drops to fourteen hours and the ROI changes.
The firms we work with that model this up front know exactly what they’re paying per task, how much time they’re saving, and whether the ROI justifies the deployment. The firms that don’t often discover the problem when the operating cost shows up on the invoice six months later.
If you want a structured way to model this for your firm, we’ve built a worksheet that walks through the token estimation, frequency modeling, and ROI calculation step by step. You can grab it here: Deploy Your First Business Agent. It’s a practical tool we use with clients during the audit phase, and it’ll give you the numbers you need to make a decision.
What Sustainable Agent Deployment Looks Like in Practice
The consulting firms that deploy agents successfully treat operating costs as a design constraint, not an afterthought. They optimize the agent to minimize unnecessary API calls, they measure both cost and time saved per task, and they iterate based on real usage data.
Here’s what that looks like in practice. A mid-sized strategy consulting firm we worked with wanted to deploy a proposal generation agent to reduce the time partners spent writing pitch decks. They modeled the operating cost at $1.20 per proposal based on their document library size and the expected token count. They ran a pilot with ten proposals, measured the time saved (twenty-two hours per proposal on average), and calculated the ROI at 180 to 1.
They deployed the agent firm-wide. Three months in, they noticed the operating cost had crept up to $2.40 per proposal. They dug into the logs and found the agent was loading the entire proposal library into context every time, even though only three to five past proposals were relevant to each new opportunity. They added a retrieval filter that pulled only the relevant documents. The operating cost dropped back to $1.20, and the output quality improved because the agent had less noise to filter.
That’s sustainable deployment. Model the cost, measure the results, iterate based on data. The firms that do this get ROI that compounds over time. The firms that don’t often abandon the agent after six months because the operating cost doesn’t match the value.
If you’re evaluating agents for your firm and you want to see what sustainable deployment looks like, See Omni for consulting firms. We’ll walk through the cost model, the workflow integration, and the iteration plan in sixty minutes. You’ll leave with a clear picture of what works and what doesn’t.
The Revenue Reality: $80K to $300K in Annual Leakage
Consulting firms doing $1M to $25M in revenue typically lose $80,000 to $300,000 annually to the three workflows we’ve covered in this article. Proposal time, research duplication, and knowledge management debt all compound across the firm, and the cost shows up as opportunity cost, not line-item expense.
A firm with six partners spending thirty hours each per quarter on proposals is losing 720 hours a year to proposal prep. At $200 per hour, that’s $144,000 in capacity that could be spent on billable client work or new business development. A firm running forty engagements a year with twenty hours of duplicated research per engagement is losing 800 hours annually, or $120,000 in consultant time. A firm where consultants spend two hours per week searching for past work is losing 100 hours per consultant per year, which compounds quickly across a team of ten or twenty.
The firms that deploy agents for these workflows recover that capacity and redeploy it into revenue-generating work. The firms that don’t continue to pay the cost every quarter.
If you want to see where the leakage is in your firm and what an agent-driven solution would look like, Book my Omni Audit. It’s sixty minutes, no deck, and you’ll leave with three outputs: a cost model for the workflows we’ve discussed, a deployment plan for the highest-ROI agent, and a clear answer on whether the operating costs justify the investment.
Where to Start
If you’re a consulting firm owner or partner evaluating AI agents, start with one repeatable workflow where the time cost is measurable and the output is consistent. Model the operating cost per task, estimate the time saved, and calculate the ROI before you deploy.
The firms that do this get agents that pay for themselves in the first quarter and compound value over time. The firms that don’t often discover the operating cost problem six months in, when the invoice arrives and the ROI doesn’t match the promise.
We’ve built agents for consulting firms across proposal generation, research synthesis, and knowledge management. The firms that get ROI right model the cost up front, optimize the agent to minimize unnecessary API calls, and measure both cost and time saved per task. If you want to see what that looks like for your firm, the Omni Audit is the next step. Sixty minutes, three outputs, no deck.
For more on how consulting firms are deploying AI agents across their operations, explore the insights library or dive into the Omni platform to see what agent-driven consulting delivery looks like in practice.