Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Thought leadership & research. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

Key Findings

Writer says Palmyra X6 can cut agent costs by 52%. Here’s how consulting firms should benchmark research and proposal workloads.

Benchmark Palmyra X6 for Lower Agent Costs
Insight ai

Benchmark Palmyra X6 for Lower Agent Costs

Sam McKay

A lower token bill matters when the work repeats

Consulting firms don’t usually have an AI cost problem because somebody used a chatbot to improve an email.

They have one when agents begin doing real work at volume.

Think about a strategy firm producing 15 proposals a month. Each proposal needs market context, relevant past work, a point of view, team bios, case studies, pricing logic, and a tailored scope. Add a handful of active engagements where teams are gathering secondary research, reviewing documents, and summarising calls. Tokens can move from an experimental line item to a recurring operating cost quickly.

Writer’s reported claim that its Palmyra X6 model can reduce AI agent costs by 52% is worth paying attention to in that context. The VentureBeat report is a vendor-backed performance claim, not a promise that every consulting firm will get the same result. Your workloads, prompts, document volumes, model routing, and quality standards all matter.

Still, the core business question is sound.

If your firm runs high-volume routine research, proposal drafting, or internal knowledge retrieval through an expensive model, you should benchmark Palmyra X6 against your current provider this quarter. Not as a model beauty contest. As a margin exercise.

For consulting and advisory firms between $1 million and $25 million in revenue, we usually see annual process leakage across proposal work, research duplication, and knowledge management land in the $80,000 to $300,000 range. AI won’t remove all of that. It can reduce a meaningful slice if you apply it to a repeatable workflow with clear review points.

The model cost is only one part of the equation. The bigger question is what a lower-cost model allows you to automate reliably without asking partners to spend their evenings correcting poor drafts.

Where token spending builds up in a consulting firm

Most partners can point to a high AI bill. Fewer can identify the workflow creating it.

That distinction matters. A $2,000 monthly bill may be acceptable if it saves 100 hours of analyst work. It may be wasteful if agents are repeatedly searching the same sources, re-reading the same internal decks, or producing lengthy output nobody uses.

There are three common patterns behind rising token consumption.

Proposals start from a blank page

A major proposal often takes 20 to 40 hours of senior and mid-level time. The work isn’t limited to writing. Someone has to locate the last relevant proposal, find a credible case study, check the latest credentials deck, pull pricing guidance, understand the client brief, and decide what should be left out.

In many firms, this work is spread across Slack, SharePoint, Google Drive, old email threads, and individual laptops. A partner may know the best prior example exists, but not where it is.

A generic AI tool can draft words, but it can’t solve that retrieval problem unless it has secure access to the firm’s approved knowledge base. When it doesn’t, people compensate by pasting large documents into prompts. That drives tokens up and creates risks around client confidentiality.

The result is a costly combination. Senior people still do the hard thinking, junior people spend hours finding inputs, and the model is asked to process more material than it needs.

Research gets repeated engagement after engagement

Every new engagement begins with some version of the same request.

Who are the major players? What has changed in the sector? What are the relevant regulations? How does the client compare with peers? Which market signals are credible enough to include in a board paper?

The first time this work is done, it’s part of the engagement. The fifth time, it is often hidden rework.

One trades-business owner in our network describes this as paying for the same insight twice. In consulting, it can be five or ten times. A team researching a regional healthcare provider might uncover useful information about procurement trends, workforce pressure, operating models, and competitor moves. Six months later, another team begins a similar assignment and starts from scratch because those findings live in a finished slide deck.

AI agents can make this worse if every engagement triggers broad web research with no structure, no source rules, and no reuse layer. You get a large token bill, a long answer, and no durable asset for the next team.

Knowledge lives inside completed work

The third pattern is knowledge management debt.

Every project generates valuable material. There are workshop transcripts, diagnostic findings, interview summaries, client-ready slides, workplans, commercial terms, and lessons from delivery. Most firms have plenty of intellectual property. They just don’t have a dependable way to find and apply it at the moment a team needs it.

A model with a high context window can help, but a larger context window isn’t a knowledge strategy. Loading everything into every prompt is expensive and often produces vague answers. A better system retrieves a small set of relevant, approved materials, cites them, and sends uncertain responses to a human reviewer.

This is where the architecture behind the agent matters more than the logo on the model.

What a proper Palmyra X6 benchmark looks like

Don’t move a production workflow because a vendor claims a percentage improvement. Build a controlled comparison around work your team already does.

Start with one workflow. Proposal preparation is often the best candidate because it has a clear input, a visible output, and a measurable amount of human effort. Research briefing is another strong option because teams can assess the quality and traceability of sources quickly.

Choose 10 to 20 representative jobs completed in the last six months. Do not cherry-pick easy requests. Include the messy ones with incomplete client briefs, specialised sectors, multiple internal references, and conflicting source material.

For each job, define:

  • The input package, such as client brief, prior proposals, case studies, approved credentials, and pricing guardrails.
  • The expected output, such as a three-page proposal draft or one-page research brief.
  • The required review standard, including factual accuracy, source citations, tone, confidentiality, and commercial fit.
  • The current model and workflow cost.
  • The time spent by analysts, managers, and partners to get to an acceptable result.

Then run the same jobs through Palmyra X6, using the same retrieval logic and prompt structure wherever possible. If one model receives better source material or more detailed instructions, you aren’t testing models. You’re testing different systems.

Measure four things.

Cost per accepted output. Token price is useful, but accepted output cost is more useful. Include model usage, orchestration, document retrieval, and the human time required to fix the draft.

Quality at first review. Ask reviewers to score factual support, usefulness, firm voice, completeness, and required revisions. A cheaper draft that needs 90 minutes of rewriting is not cheaper.

Latency. For internal research, a few extra seconds may not matter. For a proposal team working against a deadline, it can. Track the full process time, not only model response time.

Reliability. Watch for unsupported claims, missing caveats, incorrect citations, and inconsistency across runs. Consulting work carries reputational risk. A system should make reviewers faster, not give them a new quality-control burden.

If Palmyra X6 delivers a lower accepted-output cost while maintaining your standard, route suitable routine work to it. If it does not, you have still gained a useful baseline for managing AI spend.

This is the type of operational question we cover through Omni Ops. The objective isn’t to nominate one model as the winner forever. It’s to design an agent workflow where model choice can change as pricing, capabilities, and workload requirements change.

Build the agent before scaling the model bill

A well-designed agent gives the model a narrow job, the right context, and a defined handoff to a person.

Take a Proposal Generation Agent. At Enterprise DNA, our Omni ops approach uses a Proposal Generation Agent to pull past proposals, relevant case studies, approved service descriptions, and pricing guidance into a tailored first draft for a new opportunity.

The end-to-end workflow can look like this:

  1. A manager submits the prospect brief and answers six to ten structured questions about the buyer, scope, timeline, and commercial objective.

  2. The agent searches only approved internal sources. It retrieves the most relevant examples based on sector, service line, deal size, and geography.

  3. It identifies gaps. If the brief doesn’t include a required assumption or proof point, it flags that gap rather than inventing an answer.

  4. It creates a proposal structure using the firm’s preferred format, then drafts sections with links back to source material.

  5. A manager reviews the brief and draft. The partner focuses on the point of view, scope, commercial judgment, and client relationship.

  6. The final version, source selections, and edits are saved so the next proposal gets better inputs.

The model may be Palmyra X6 for early drafting and routing, another model for more complex reasoning, or a combination. What matters is that each task has a cost and quality target.

The same design applies to the Research Agent. It runs structured industry and company research at the start of every engagement, producing sources, summaries, and a one-page brief. It should have clear rules about acceptable sources, dates, industry-specific terminology, and how to distinguish fact from inference.

A good Research Agent does not deliver 12 pages of generic market commentary. It gives a project team a concise starting point, a source trail, and a list of questions that require primary research or client validation.

Finally, the Knowledge Agent reads the decks, documents, and meeting transcripts your firm produces, then answers questions across that approved corpus. That means a team can ask, “What have we learned about operating model redesign in mid-market manufacturers?” and get a useful response supported by internal references.

This is not a replacement for expert judgment. It’s a way to stop highly paid people spending 45 minutes trying to remember where a relevant slide lives.

For a broader view of how these systems fit together, review the Omni platform and the practical material in our AI guides.

The cost case is bigger than token pricing

A 52% reduction in agent model costs sounds compelling. But model savings alone may not justify a migration effort if your overall usage is modest.

The stronger case comes from combining cost control with workflow improvement.

Say your firm spends $3,000 a month on model usage for research and proposal agents. A 52% reduction, if it holds in your benchmark, is meaningful. It could free up roughly $18,000 a year before considering any changes in volume.

Now consider the human side. If a Proposal Generation Agent helps reduce 10 hours of preparation time across eight significant proposals each month, that is 80 hours returned to the firm. Not every hour becomes immediate cash savings. Some gets redeployed to client work, business development, or better delivery. But it gives you capacity without automatically adding headcount.

The same applies to research. A Research Agent that consistently saves two to four hours at the start of an engagement can make a difference across 30 or 50 projects a year. The value depends on your billing model, utilisation, and ability to redeploy the time.

The most practical way to assess this is to map the current process first. See Omni for consulting firms if you want the process and value questions framed around consulting work rather than generic AI use cases.

If you want to identify the first agent, its source systems, review points, likely effort reduction, and model cost profile, Book a 60-min Omni Audit. It is a working session, not a sales deck.

Keep the benchmark grounded in commercial reality

There are a few traps consulting firms should avoid.

First, don’t benchmark using a single ideal prompt. Production work is variable. A useful test includes rushed proposal requests, thin briefs, difficult client terminology, and incomplete internal documentation.

Second, don’t treat output length as output quality. Agents can use lots of tokens to create polished text that says very little. Require citations, decision-relevant findings, and structured output that fits into your delivery process.

Third, don’t give every agent access to every document. Client confidentiality, commercial sensitivity, and quality all improve when access is scoped to the task. Your retrieval layer should enforce permissions, retention rules, and source approval.

Fourth, don’t put partners in the role of full-time AI editors. The workflow should direct uncertain decisions to the appropriate person. A partner should review a proposal’s commercial position, not spend 30 minutes correcting invented company facts.

There is also a strategic point here. The lowest-cost model may be the right choice for high-volume classification, extraction, first-draft research, and standard summaries. It may not be the right choice for every executive narrative or complex advisory judgment. Model routing lets you make that distinction deliberately.

You don’t need to bet the firm on one provider. You need a process that makes it easy to compare providers and assign work to the right model at the right cost.

A practical first step for this quarter

Start by pulling your last 90 days of AI usage and identifying the top three recurring tasks. Don’t begin with departments. Begin with work units.

You might find that 40% of spend relates to proposal drafting, 30% relates to broad research requests, and the rest sits in ad hoc experimentation. Or you may discover that the real cost is not tokens at all. It is managers manually preparing documents for agents because the firm’s knowledge is poorly organised.

Either way, the answer is clearer after a short audit.

If you need a simple way to prepare internally, our Deploy Your First Business Agent resource is a practical worksheet for choosing a workflow, defining inputs and outputs, assigning a human reviewer, and estimating the value of the time saved. You can also download the worksheet directly and use it with your leadership team before committing to a pilot.

Palmyra X6 may prove to be a lower-cost option for your routine agent workloads. Benchmarking will tell you. The larger opportunity is to build a repeatable operating model where proposals, research, and hard-won knowledge stop being recreated from scratch.

For that, the AI audit for consulting firms is a sensible place to start. In 60 minutes, we map the workflow, identify the best first agent, and outline the operational and commercial case. There is no deck to sit through.

When you’re ready to turn the benchmark into a practical plan, Book my Omni Audit.