Your AI Costs Are Becoming a Client Line Item
A partner I spoke with recently told me their firm’s AI tooling bill had tripled in six months. Nobody had signed off on that. It just happened, quietly, as every consultant on staff started leaning on AI-assisted research and drafting to move faster on client work. Nobody was tracking which engagements were burning the tokens, or why one project cost four times another to produce the same type of deliverable.
That’s the shift happening right now across advisory firms. AI-assisted research and report generation stopped being a productivity nice-to-have and started becoming a real cost of goods sold. Every prompt, every research pass, every draft revision has a token cost attached to it, and those costs roll up per project the same way staff hours do. The difference is almost nobody is tracking it that way yet.
The Network Dispatch’s August 2026 note on this made a point worth taking seriously: clients are going to start asking. Not in a hostile way, necessarily, but the same way they’ve always asked about staffing mix and realization rates. If a firm is using AI to generate a chunk of a deliverable, sophisticated clients will want to know what that costs and whether they’re paying appropriately for it. Firms that can’t answer that question with numbers are going to look sloppy at best, and get re-negotiated at worst.
Why this matters more for firms your size
If you’re running a $1M-$25M consulting or advisory practice, you don’t have a finance team dedicated to parsing API usage logs against project codes. You’ve got a few partners, a handful of senior consultants, and probably a mix of AI tools that different people adopted on their own initiative because they worked well enough to keep using.
That’s exactly the setup where token cost sneaks past you. Someone on a research-heavy engagement runs dozens of iterative queries trying to get a synthesis right. Someone else drafts three full proposal versions before landing on the one that goes out. Multiply that across a dozen active engagements and you’ve got real dollars moving through a channel nobody’s watching.
We see this pattern show up in the same leakage band we track across advisory firms generally, somewhere between $80,000 and $300,000 a year in unmanaged AI and process cost once you count the token spend, the redundant research hours, and the rework that comes from nobody having a system for any of it. That’s not a scare number. It’s the kind of range that’s typical for firms of this size once you actually go looking.
The three places the cost actually hides
Proposal and pitch work. Senior people are still writing decks and proposals from scratch every time a new opportunity comes in. Your win rate might be perfectly healthy, but the cost of getting to that win is brutal, often 20 to 40 hours of senior time on a major proposal. Now layer AI drafting into that process without a structured way to reuse past work, and you’re paying token costs on top of the hours, with no guarantee the AI-assisted draft is actually faster than what a well-organized template would have produced.
Research and synthesis. Every engagement kicks off with weeks of secondary research, industry scans, competitor reviews, market sizing. Most of that work gets repeated project to project because there’s no shared memory of what the firm already knows. AI tools make this research faster to produce, but faster production of duplicate work is still duplicate work. It just costs tokens now instead of only hours.
Knowledge management debt. Every project your firm runs produces real intellectual property. A framework, a data model, a client-specific insight that would apply somewhere else. Almost none of it makes it back into a form the next team can use. So the firm pays for the same insight twice, once when it’s created and again when a different team reinvents it six months later, this time with an AI tool doing the reinventing at a token cost that nobody’s logging.
None of these are new problems. What’s new is that AI has made the first two faster without making them cheaper in a way anyone’s tracking, and it’s made the third one look solved when it isn’t.
What tracking this actually looks like
This isn’t about installing a dashboard and hoping people check it. It’s about building the cost visibility into the actual workflow, at the point where the work happens, using agents that do the work and report what it cost as a byproduct.
We build this inside Omni as a set of named agents that handle specific pieces of the workflow, not a general chatbot bolted onto your existing process.
The Research Agent runs structured industry and company research at the start of every engagement. It pulls from defined sources, produces a one-page brief with citations, and logs what that research actually cost in tokens against the project code. Your consultants get a usable starting brief in hours instead of weeks, and you get a real number to compare against last quarter’s research spend on a similar engagement.
The Proposal Generation Agent pulls from your firm’s past proposals, case studies, and pricing history to produce a tailored first draft for a new opportunity. Instead of a senior partner starting from a blank document, they’re editing a draft that already reflects how your firm actually writes and prices this kind of work. The cost of producing that draft is visible per proposal, so you can finally answer the question of whether your AI-assisted proposal process is actually cheaper than the old way, not just faster-feeling.
The Knowledge Agent reads every deck, document, and meeting transcript your firm produces and answers questions across that whole corpus. This is the piece that turns your project archive from a graveyard into an asset. A consultant starting a new engagement can ask what the firm has already learned about a sector or a client type, and get a synthesized answer pulled from real prior work, instead of pinging six people on Slack and hoping someone remembers.
Each of these agents does real work and produces a real deliverable. The token cost tracking isn’t a separate reporting layer, it’s a natural output of routing the work through a system built to log it. That’s the difference between “we think AI is saving us time” and “here’s what this proposal cost to produce, broken down by hours and tokens.”
What this means for how you price and staff
Once you can see token cost per deliverable, a few things change fast. You can price engagements that lean heavily on AI-assisted research differently than ones that don’t, because you actually know the cost delta. You can stop overstaffing research phases with senior people whose time is worth more than the task requires. And when a client asks what they’re paying for, you have an answer that’s specific instead of defensive.
This also protects you on the other side. Firms that get caught flat-footed on this question, unable to explain their AI cost structure, are the ones that end up in fee negotiations they didn’t see coming. Firms that get ahead of it are the ones who can say, credibly, that their AI-assisted process delivers more value per dollar than the traditional alternative. That’s a stronger position to negotiate from, not a weaker one.
If you want a structured way to start building this into how your team works, our Deploy Your First Business Agent guide walks through picking one workflow, wiring up an agent to handle it, and measuring the before-and-after in real numbers. You can grab the direct download here if you want to work through it with your team this week. It’s built as a practical worksheet, not a theory document, and it’s a reasonable first step if you’re not ready for a full audit yet.
For a broader look at how these agents fit together across a firm’s operations, our Omni Ops page walks through the agent categories we build most often, and our insights section has more on how advisory firms specifically are adapting their delivery models as AI-assisted work becomes standard rather than novel.
Where the Omni Audit fits
Reading about this is useful. Knowing your own numbers is what actually changes anything. That’s what the Omni Audit is for.
It’s a 60-minute session, no deck, no sales pitch dressed up as a workshop. We walk through your actual proposal process, your research workflow, and how project knowledge moves (or doesn’t) across your firm. You walk away with three concrete outputs, a map of where your token and process costs are actually going, a specific estimate of the annual leakage in your business, and a short list of which agent would save you the most, starting with either the Research Agent or the Proposal Generation Agent depending on where your bottleneck sits.
If you want to see how this works specifically for firms like yours, see Omni for consulting firms walks through the audit process and what past sessions have surfaced for firms your size. It’s worth ten minutes even if you’re not ready to book yet.
When you are ready, Book a 60-min Omni Audit and bring your last three proposals and a rough sense of your AI tooling spend. We’ll build the picture from there.
The real question to ask this week
Before your next client engagement kicks off, ask your team a simple question. What did the research phase cost, in hours and in AI spend, on the last three projects that looked similar? If nobody can answer that with a number, you’ve found your starting point.
This isn’t a problem you solve by cutting AI use. It’s a problem you solve by making the cost visible, project by project, the same way you already track billable hours. Firms that build that habit now are the ones who’ll be able to answer the client’s question calmly when it comes. And based on what we’re seeing across the market, it’s coming for most firms within the next year or two, not the next decade.
If you want a second set of eyes on where your firm actually stands, our guides section has more detail on agent-based workflows for advisory firms, and the Omni Advisory page covers how we scope this kind of work for firms in the $1M-$25M range specifically. Or skip straight to the source and see Omni for consulting firms to get a feel for what the audit actually covers.
The token cost conversation is coming to your client meetings whether you’re ready for it or not. Better to walk in with the numbers already in hand.
If you’d rather just talk it through, Book my Omni Audit and we’ll spend the hour on your firm’s actual numbers, not a generic framework.