Why Cheaper AI Models Won't Cut Your Agent Costs
The promise sounded straightforward: AI model prices drop 90%, so enterprise costs should follow. Consulting firms heard that pitch and started deploying agents to handle research, proposal drafting, and knowledge synthesis. Then the bills arrived, and the math didn’t match the marketing.
A recent industry report confirms what we’re seeing across mid-market firms. Cheaper models don’t translate to cheaper operations when agent workflows consume thousands of queries per deliverable. One research agent might call a model 300 times to compile a single client brief. A proposal agent pulls context from dozens of past projects, rewrites sections, checks compliance, and iterates on feedback. Each step burns tokens. The per-query cost is tiny, but the volume compounds fast.
For consulting firms running on tight margins, this matters. You’re not paying for a chatbot that answers three questions a day. You’re paying for agents that complete work, and completed work in a knowledge business requires context, iteration, and synthesis. That’s where the token count explodes, and where firms that budgeted on per-query pricing hit a wall.
The Real Cost Structure of Agent Work
Most consulting firms think about AI costs the way they think about SaaS: a monthly seat fee, maybe some usage tiers. That model breaks when you deploy agents that do actual deliverable work. Here’s why.
A senior consultant writing a proposal pulls from memory, past projects, client conversations, and industry knowledge. They draft, revise, check for consistency, and tailor the language to the buyer. That process might take 20 hours. An agent doing the same work doesn’t take 20 hours, but it does make hundreds of API calls. It reads past proposals to find relevant case studies. It pulls pricing data from your CRM. It checks the latest version of your service descriptions. It drafts three versions of the executive summary and picks the best one based on client tone. Each of those steps is a query, and each query costs money.
The model price per token dropped, but the number of tokens per completed task didn’t. In fact, it went up. Agents are more capable now, so firms ask them to do more complex work. A research brief that used to be a bulleted list is now a formatted document with citations, competitive analysis, and a risk assessment. The output is better, but the token budget is five times what it was six months ago.
We see this with firms running our Omni Ops agents. A Proposal Generation Agent might consume 50,000 tokens to produce one tailored pitch deck. That’s cheap per token, but it’s 50,000 tokens you didn’t budget for when you looked at the per-query pricing page. Multiply that by 30 proposals a month, and you’re spending real money even with the cheapest models on the market.
Why Per-Query Pricing Misleads Consulting Firms
The per-query pricing model works fine for simple use cases. A customer service bot that answers “Where’s my order?” doesn’t need much context. It reads a tracking number, calls an API, and returns a result. Total token count: maybe 200.
Consulting work is different. Every deliverable requires context. A Research Agent starting a new engagement doesn’t just answer one question. It reads the client’s website, pulls industry reports, summarizes competitor positioning, and flags regulatory risks. That’s not one query. That’s a workflow with 15 steps, each pulling data from different sources, each requiring the agent to synthesize and format the output.
One firm we work with deployed a Knowledge Agent to answer questions about past projects. The agent reads every proposal, deck, and meeting transcript the firm has produced in the last five years. When a partner asks “What did we recommend for supply chain optimization in the pharma sector?”, the agent doesn’t just search for keywords. It reads relevant documents, extracts the recommendations, checks for contradictions across projects, and summarizes the findings. That single question might trigger 100 API calls.
The per-query model assumes each query is independent. Agent workflows assume each query builds on the last. The cost difference is enormous, and most firms don’t realize it until they’ve been running agents for a month and the bill is three times what they expected.
If you’re evaluating agent costs for your firm, the AI audit for consulting firms walks through the real token economics of your specific workflows. We map your deliverables to agent tasks and show you the actual usage patterns, not the theoretical per-query math.
Budget on Deliverables, Not Queries
The firms that control agent costs don’t track queries. They track completed work. How many proposals did the agent draft this month? How many research briefs? How many client questions did the Knowledge Agent answer with a full synthesis instead of a one-line response?
That shift in thinking changes how you budget. Instead of estimating queries per day, you estimate deliverables per month. A mid-market consulting firm might produce 25 proposals, 40 research briefs, and 100 knowledge lookups in a month. Each of those has a token budget based on real usage, not a guess about how many times someone will click a button.
We built our Proposal Generation Agent with this model in mind. It doesn’t optimize for the fewest queries. It optimizes for the best output, which means pulling more context, running more comparisons, and iterating more drafts. The token count is higher, but the output is good enough that partners send it to clients with minimal edits. That’s the trade-off: higher per-deliverable cost, but the deliverable actually ships.
One trades-business owner in our network describes it this way: “I don’t care if the agent makes 500 API calls to write a proposal. I care that the proposal is 90% done when I open it, and I can send it the same day. The old way cost me 20 hours of senior time. The new way costs me $30 in API fees and two hours of review. That’s not even close.”
The math works when you price it correctly. A senior consultant billing $250 an hour who spends 20 hours on a proposal costs the firm $5,000 in opportunity cost. An agent that costs $50 in API fees to produce the same proposal is a bargain, even if it’s using the most expensive model on the market. The problem is that firms see the $50 and panic because they budgeted $5 based on per-query pricing.
The Three Cost Traps Consulting Firms Hit
The first trap is underestimating context requirements. Consulting deliverables aren’t standalone. A proposal references past work, pricing history, and client conversations. An agent needs access to all of that, which means reading documents, pulling CRM data, and cross-referencing notes. Each context pull is a query, and the context window fills up fast.
The second trap is iteration. Agents don’t produce perfect output on the first try. A research brief might go through three drafts before it’s client-ready. Each draft is a new set of queries, and each revision pulls fresh context to incorporate feedback. Firms that budget for one pass per deliverable end up paying for three.
The third trap is scope creep. Once you have an agent that writes proposals, people start asking it to do more. Can it write the follow-up email? Can it draft the statement of work? Can it summarize the kickoff call and update the project plan? Each of those requests is reasonable, and each one adds token usage that wasn’t in the original budget.
We see this most often with Knowledge Agents. A firm deploys one to answer questions about past projects. Six months later, it’s being used to onboard new hires, draft thought leadership, and prep partners for sales calls. The usage is 10x the original estimate, not because the agent is inefficient, but because it’s useful and people find new ways to use it.
The solution isn’t to lock down access. It’s to budget for growth. Assume your agent usage will double in the first six months. Assume people will find creative ways to use agents you didn’t plan for. Build that into your cost model from the start, and you won’t be surprised when the bill climbs.
For a practical breakdown of how to map your current workflows to agent tasks and estimate real token usage, we’ve put together a worksheet that walks through the process step by step. You can grab it here: Deploy Your First Business Agent. It’s the same framework we use when we run an audit, and it’ll give you a realistic picture of what your agent costs will look like once you’re in production.
What Agent Economics Look Like in Practice
Let’s walk through a real example. A 15-person consulting firm wants to deploy a Research Agent to handle the secondary research that kicks off every client engagement. Right now, a junior consultant spends two weeks per project reading industry reports, summarizing competitor moves, and building a briefing document. The firm runs 20 projects a year, so that’s 40 weeks of junior time, or roughly $80,000 in fully loaded cost.
The Research Agent can do the same work in 24 hours. It reads the same reports, pulls the same data, and produces a formatted brief with citations. The output isn’t perfect, but it’s 80% of the way there, and a senior consultant can review and finalize it in four hours instead of starting from scratch.
Here’s the token math. The agent reads 50 documents per project, averaging 5,000 tokens each. That’s 250,000 tokens of input. It writes a 10-page brief, which is about 8,000 tokens of output. Add in the intermediate steps where it summarizes, compares, and formats, and you’re looking at 300,000 tokens per project. At current pricing for a mid-tier model, that’s about $6 per project, or $120 a year for 20 projects.
The cost savings are obvious: $80,000 in labor vs. $120 in API fees. But the real win is speed. The firm can now kick off a project the day the contract is signed instead of waiting two weeks for research. That compresses the sales-to-delivery cycle, which means faster cash collection and happier clients.
The catch is that the firm didn’t budget $120. They budgeted $20 because they looked at per-query pricing and assumed each project would be a handful of questions. When the first month’s bill came in at $60 instead of $10, they thought something was broken. Nothing was broken. The agent was doing exactly what it was supposed to do, and the token usage was in line with the complexity of the task.
How to Model Agent Costs Before You Deploy
Start with your deliverables. List every recurring output your firm produces: proposals, research briefs, client reports, onboarding documents, knowledge base articles. For each one, estimate how many you produce per month and how much senior time each one currently takes.
Next, map each deliverable to an agent workflow. A proposal might require the agent to read five past proposals, pull pricing from your CRM, check your current service descriptions, and draft three sections. A research brief might require reading 30 industry documents, summarizing competitor positioning, and formatting a 10-page output. Write down each step in the workflow and estimate the token count for each step.
You don’t need to be precise. Use ranges. A document read is 2,000 to 10,000 tokens depending on length. A summary is 500 to 2,000 tokens. A draft section is 1,000 to 5,000 tokens. Add it up, multiply by the number of deliverables per month, and you’ve got a realistic usage estimate.
Then stress-test it. Assume usage doubles in six months. Assume people find new ways to use the agent. Assume the agent needs three passes per deliverable instead of one. If the cost is still a fraction of your current labor cost, you’re in good shape. If it’s not, you need to rethink the workflow or the pricing model.
The firms that get this right treat agent costs like they treat contractor costs. You wouldn’t hire a contractor without knowing the scope, the rate, and the expected hours. Don’t deploy an agent without knowing the deliverables, the token budget, and the expected usage. The math is different, but the discipline is the same.
We run this analysis as part of the Omni Audit for consulting firms. It’s a 60-minute working session where we map your workflows, estimate token usage, and show you what your agent costs will look like in production. No deck, no sales pitch. Just the numbers and a plan. Book a 60-min Omni Audit and we’ll walk through it together.
The Agents That Actually Pay Off
Not every agent is worth deploying. Some tasks are too simple to justify the setup cost. Some are too complex to trust to an agent without heavy oversight. The agents that pay off are the ones that handle high-volume, high-context work that currently burns senior time.
Proposal Generation Agents are the most common win. Every consulting firm writes proposals, and every proposal pulls from the same pool of past work, case studies, and pricing. An agent can draft a tailored proposal in minutes instead of hours, and the output is good enough that partners can review and send it the same day. The token cost is higher than a simple chatbot, but the time savings are massive.
Research Agents are the second most common win. Every engagement starts with research, and most of that research is repeated across clients. An agent can read industry reports, summarize competitor moves, and flag risks faster than a junior consultant, and it doesn’t need to be trained on your firm’s methodology. The output needs review, but it’s a starting point instead of a blank page.
Knowledge Agents are the third win, but they’re harder to deploy. A Knowledge Agent needs access to your entire corpus of past work, which means cleaning up file structures, tagging documents, and building a search layer that actually works. The payoff is huge once it’s running, but the setup cost is real. We usually recommend starting with a Proposal or Research Agent and adding a Knowledge Agent once you’ve got the infrastructure in place.
For more on how these agents fit into a broader automation strategy, the EDNA insights library has case studies and workflow breakdowns from firms that have deployed them in production.
Why the Omni Audit Matters Now
Cheaper models didn’t make agent costs disappear. They made agent costs harder to predict. Firms that deploy agents without understanding token economics end up either overpaying or underutilizing. The Omni Audit fixes that.
It’s a 60-minute working session where we map your workflows, estimate token usage, and show you what your agent costs will look like in production. You walk away with three outputs: a workflow map, a token budget, and a deployment plan. No deck, no sales pitch. Just the numbers and a next step.
The firms that run the audit before they deploy save months of trial and error. They know which agents to build first, which workflows to automate, and what their costs will look like at scale. The firms that skip the audit end up rebuilding agents, rewriting prompts, and renegotiating contracts because the math didn’t work out.
If you’re running a consulting firm and you’re thinking about deploying agents, don’t guess on the economics. Book your Omni Audit and we’ll walk through it together. You’ll know exactly what it costs, what it saves, and whether it’s worth doing before you write a line of code.
Cheaper models are a tool, not a solution. The firms that win are the ones that understand the economics, budget for the work, and deploy agents that actually ship deliverables. That’s what we build at Omni, and that’s what the audit is designed to show you.