Agent Costs Don’t Scale Like Software
The scale assumption that can hurt consulting firms
A lot of consulting partners see agentic AI through a software lens.
Build the workflow once. Run it 100 times. Watch the unit cost fall.
That logic works in parts of software. It doesn’t always work in multi-step AI workflows.
A research agent that gathers public sources, reads annual reports, searches internal project files, drafts a market brief, checks citations, and sends exceptions to a consultant has costs at every step. More work may create better process data and expose opportunities to improve prompts or routing. It does not automatically make each completed brief cheaper.
The warning reported in the Computer Weekly coverage of Gartner’s view on agentic AI matters for consulting and advisory firms because your cost base isn’t just model usage. It includes senior review, specialist validation, data access, exception handling, and the commercial cost of getting an answer wrong.
If a consulting firm treats an agent like a cheap software licence, it can scale a workflow that creates more work for its best people.
The right question isn’t, “Can we build an agent?”
It is, “What does a completed, trustworthy task cost us, and what part of that cost actually falls as volume rises?”
That is the question we work through in the AI audit for consulting firms. For firms between $1 million and $25 million in revenue, the difference is material. We commonly see $80,000 to $300,000 a year tied up in repeated research, proposal production, knowledge retrieval, and review work that is not visible in a normal P&L line.
Why agent workflows don’t behave like SaaS
A conventional software product has a large fixed cost and a low incremental cost. Once the product is built, each new user often adds only a small amount of hosting and support cost.
An agentic workflow has some fixed costs too. You need to map the process, connect systems, define permissions, write instructions, test output, and create escalation paths. But its variable costs can remain meaningful because each task requires active computation and, in many cases, human judgement.
For a consulting workflow, the cost per task usually has six parts.
-
Model usage. Every research pass, document summary, analysis step, draft, revision, and quality check consumes model capacity.
-
Tool and data usage. Search APIs, company databases, transcription, document extraction, CRM lookups, and specialist data sources may charge by request, document, seat, or volume.
-
Retrieval and context preparation. Pulling past case studies, pricing schedules, project documents, and meeting notes into an answer takes processing. Poorly structured source material can make this expensive and unreliable.
-
Workflow orchestration. An agent that decides which source to check next, asks another agent to validate a claim, or retries a failed request creates more workflow steps. Those steps are useful only if they improve the final result enough to justify their cost.
-
Human validation. This is the one leaders routinely undercount. A principal checking a client-facing recommendation is not a free safety net. Their time has a real delivery cost and an opportunity cost.
-
Exceptions and remediation. Every firm has awkward files, missing source documents, vague client requests, stale pricing, and conflicting information. These are not edge cases. They are daily operating conditions.
Volume can lower certain elements. Your team gets better at handling exceptions. You may negotiate better vendor pricing. A knowledge base becomes more complete. Reusable workflow components reduce build time for the next use case.
Still, more tasks can also mean more stale data to manage, more quality checks, more sensitive client information, and more cases that don’t fit the happy path.
That is why the unit of analysis should be a completed task, not a monthly platform bill.
Start with the work, not the agent label
A consulting firm doesn’t need an “AI strategy” before it can make a sensible decision. It needs to identify one repeated task with a clear output, a known owner, and a measurable cost.
Three areas tend to stand out.
Proposal and pitch production
A major proposal can consume 20 to 40 hours across partner, director, manager, analyst, and design time. The work is rarely just writing. Someone needs to interpret the brief, find relevant credentials, check claims, pull team biographies, shape a commercial approach, source pricing inputs, and make sure the story fits the buyer.
The win rate might be acceptable. The cost of sale is what quietly gets out of hand.
A Proposal Generation Agent in Omni ops can pull approved past proposals, case studies, capability statements, rate cards, and discovery notes into a first draft. It can flag missing information, suggest relevant proof points, and produce a structured response for a proposal lead to edit.
That does not mean you hand the agent a tender and send the output to a prospect.
The partner still owns the deal logic. The proposal lead still checks the claims. Finance still approves commercial terms. The question is how much of the assembly and retrieval work can be removed before human judgement is needed.
Measure the workflow from opportunity intake to an approved draft. Don’t just measure the five minutes it takes to generate text.
Engagement research and synthesis
Most advisory engagements start with an unglamorous period of reading. Analysts search company filings, investor material, competitor sites, industry reports, news coverage, client documents, and previous engagement material. Then they turn a large pile of information into a usable point of view.
Some of this work should be repeated. A new client in a familiar sector should benefit from what the firm already knows.
Yet many firms begin each engagement as if the prior work never existed. Research gets recreated in a different format by a different team. Sources are saved to local folders. The final deck contains insights, but not always the supporting material in a reusable form.
A Research Agent can run a defined starting workflow. It receives a client name, sector, geography, business question, and engagement context. It then gathers sources within agreed rules, creates summaries, identifies gaps, and produces a one-page brief with citations and a research log.
That is a valuable use case, but it is also a good illustration of why cost doesn’t automatically decline at scale.
A simple brief might require 10 sources and one validation pass. A strategic market-entry question may require 50 sources, several documents behind logins, internal knowledge retrieval, conflicting data checks, and partner review. The agent isn’t producing the same unit of output every time.
Your cost model must separate straightforward assignments from complex ones. If you average them together, you will either over-engineer simple work or underprice difficult work.
Knowledge management debt
Every project creates intellectual property. Client interviews, working papers, delivery decks, workshop outputs, proposal notes, research packs, and meeting transcripts all contain useful material.
Then the project closes. The knowledge goes into SharePoint, Google Drive, Teams, a project folder, a consultant’s laptop, or a folder with a name nobody will search again.
The next team spends days reconstructing an answer that the firm has already paid to develop.
A Knowledge Agent can read approved decks, documents, and meeting transcripts, classify them against a firm taxonomy, and answer questions across the corpus. It can help a manager find examples of operating model work in healthcare, previous pricing structures for a comparable client size, or recommendations that appeared in several transformation programmes.
But knowledge agents have a major validation cost. They need permission controls, source citations, clear treatment of confidential client material, retention rules, and an answer style that makes uncertainty obvious. A confident answer based on the wrong engagement is worse than no answer.
That is why firms should begin with a bounded knowledge domain. Start with approved, non-sensitive material from a defined practice area. Measure retrieval quality, answer usefulness, and the minutes saved in a live pursuit or project setting.
Build a per-task cost model before rollout
You don’t need a finance team to create a practical model. You need a process owner, a sample of real work, and the discipline to count the whole workflow.
Take one proposed agent task and document the following.
| Cost area | What to measure |
|---|---|
| Task volume | Tasks per week, month, and expected peak periods |
| Task complexity | Simple, standard, and complex cases |
| Model and tool cost | Average usage across each workflow step |
| Human touch time | Minutes for review, correction, approval, and exception handling |
| Failure rate | Tasks requiring rework, escalation, or a manual restart |
| Value created | Time saved, faster response, improved reuse, or avoided external spend |
For example, a Research Agent may save an analyst six hours on a standard engagement brief. That is not the full benefit until the firm asks what it takes to get a reliable brief.
If the workflow uses several tools, runs 15 research steps, and takes a manager 30 minutes to validate sources and sharpen the point of view, the unit cost may still be worthwhile. But it is not zero. If one in five briefs needs substantial rework because the client context was unclear, the economics change again.
The same is true for proposals. Saving 12 hours of analyst assembly time is valuable. Saving those hours while adding two hours of partner review is still likely positive. Saving 12 hours but creating a document that makes unsupported claims or uses out-of-date pricing is not positive.
The point is not to demand perfect measurement before you start. It is to avoid scaling based on a demo.
For a broader view of how these workflows fit with operating processes, review the practical material in our AI learning resources. The important shift is from asking what the tool can generate to asking what the business can reliably deliver.
The validation burden is where many plans break
Consulting firms sell judgement. Clients pay you to understand their context, challenge assumptions, and make a recommendation that holds up when it reaches an executive team.
That changes the design of an agent workflow.
A good workflow should not conceal validation. It should make validation faster and more precise.
For research, that means source links, dates, confidence markers, and a clear list of unanswered questions. For proposals, it means marking which case studies were selected, identifying where commercial inputs came from, and preventing unapproved material from being used. For knowledge retrieval, it means showing the original source document and respecting client and practice permissions.
The agent should also know when to stop.
If it cannot find credible sources, it should say so. If it finds conflicting financial figures, it should flag the conflict. If a query involves a restricted client or a sensitive commercial issue, it should route the work to a named owner.
This is not bureaucracy. It is how you stop senior people from reviewing every output from scratch.
One trades-business owner in our network describes this as “reviewing the evidence, not redoing the work.” The same principle applies in consulting. The ideal reviewer sees a concise output, checks the evidence, makes a judgement call, and moves on.
Where scale does help, and where it doesn’t
It would be wrong to say agent workflows cannot get more efficient.
They can. A firm can reduce build cost by reusing connectors, source libraries, evaluation criteria, and approval patterns. It can standardise an intake form. It can tune a workflow after seeing the same failure mode several times. It can create a stronger internal knowledge base.
The Omni platform is designed around this practical reuse. The objective is not to build a different isolated bot for every partner. It is to create agents that work inside a governed operating model.
Still, task-level costs remain.
A client-facing research brief needs a view on source quality. A proposal needs commercial approval. A knowledge answer may need a confidentiality check. More volume may improve the process, but it does not remove the need for control.
That is why a staged rollout works better than a firm-wide launch.
Start with a narrow workflow. Run a representative sample, perhaps 20 to 50 tasks. Track the baseline manual time, agent cost, human review time, exception rate, and output quality. Then decide what to improve before expanding to a second team or use case.
If you want a worksheet to structure that first deployment, download Deploy Your First Business Agent. The direct version is available here: Deploy Your First Business Agent checklist. It helps you define the task, owner, inputs, approval point, and measures before you spend months building.
A practical 60-minute test for your firm
Before approving a multi-agent programme, bring together the partner or GM who owns the result, the person doing the work now, and the person responsible for data or systems.
In one hour, answer these questions.
- Which recurring task consumes the most expensive time?
- What is the output, and what does “good” look like?
- Which steps require model or tool calls?
- Which steps still need human judgement?
- What errors would create client, commercial, or confidentiality risk?
- How many standard tasks occur each month?
- What would need to be true for the workflow to save real capacity rather than simply move work around?
If your answer is vague, don’t scale yet. Run a smaller test.
If the task is clear, the inputs are accessible, the approval point is defined, and the value is visible, you have a better foundation for an agent than most firms that start with a technology purchase.
You can Book a 60-min Omni Audit when you want to put real numbers around that decision. There is no deck. In 60 minutes, we map the manual workflow, identify the points where an agent can reduce effort without weakening quality, and outline the expected financial upside and implementation priorities.
Don’t scale the wrong unit economics
The commercial opportunity for consulting firms is real.
Proposal work can move faster without starting from a blank page. Research can become more repeatable. Existing intellectual property can become accessible at the moment a team needs it. Those improvements can release capacity, protect margin, and reduce the $80,000 to $300,000 of annual leakage that firms of this type often carry.
But the opportunity is not “replace all human work with agents.”
It is to remove the repeated assembly, retrieval, and administrative judgement that sits around the work clients actually value. That requires a cost model that includes technology, people, validation, and exceptions.
Measure the cost of a trustworthy completed task. Improve the workflow. Then scale what has proved itself.
For a consulting-specific view of where to begin, see Omni for consulting firms. If you’re ready to identify the highest-value first workflow in your own firm, Book my Omni Audit.