Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Thought leadership & research. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

Key Findings

Major AI vendors are switching from flat fees to token billing. Law firms that don't renegotiate now will lose forecast control and case profitability.

Token Pricing Will Kill Your Case Margins (Act Now)
Insight ai

Token Pricing Will Kill Your Case Margins (Act Now)

Sam McKay

You signed a Salesforce or Microsoft contract last year with a clean monthly seat price. Now your vendor rep is floating a new pricing model tied to “consumption units” or “token usage.” The email sounds friendly, but the math is brutal. What cost you $180 per user per month could triple in a busy quarter, and you won’t know until the invoice arrives.

This isn’t hypothetical. Salesforce, Microsoft, and a dozen other enterprise AI vendors are moving away from predictable seat-based pricing toward token-metered billing. For law firms, this shift is catastrophic. You can’t forecast case profitability when your software bill swings 200% based on how many discovery documents your associates ran through the system. You can’t budget headcount when a single complex matter burns through your annual AI allocation in six weeks.

The window to renegotiate is now. Once your renewal auto-converts to the new model, you’re locked in for another 12 to 36 months. By then, your competitors will have moved to purpose-built agents that cost a fraction per task and don’t punish you for using them.

Why Token Pricing Breaks Law Firm Economics

Traditional legal software charged per seat or per matter. You knew the cost before you opened a file. Token-based AI flips that. Every API call, every document ingested, every question asked to a language model burns tokens. The vendor meters usage in real time and bills you at the end of the month.

For a consumer app, tokens are invisible. For a law firm managing 40 active matters across three practice areas, tokens are a budget black hole. Here’s what happens in practice.

Your associate runs a 300-page contract through an AI review tool. The vendor’s model tokenises every word, every clause, every footnote. That single document might consume 150,000 tokens. At $0.02 per 1,000 tokens (a mid-tier rate), you just spent $3 on one contract. Multiply that by 80 discovery documents in a commercial dispute, and you’ve burned $240 before anyone billed an hour.

Now add the intake agent that answers after-hours calls. Each five-minute conversation generates a transcript, a conflict check, and a summary memo. Another 8,000 tokens per call. If you get 60 after-hours inquiries in a month, that’s 480,000 tokens, or $9.60 at the same rate. Sounds cheap until you realise your vendor charges $0.06 per 1,000 tokens for real-time voice, not $0.02. Now it’s $28.80, and that’s just one feature.

The real damage isn’t the absolute dollar amount. It’s the unpredictability. A quiet month costs you $400. A busy month with two large matters costs $1,800. You can’t staff for that variance. You can’t quote flat fees. You can’t even decide whether to use the tool without running a cost-benefit calculation every time an associate opens a file.

Most firms we work with don’t discover the problem until month three, when the first big invoice lands. By then, they’ve trained the team on the tool, integrated it into their workflow, and told clients they’re using “cutting-edge AI.” Ripping it out means retraining, workflow disruption, and a credibility hit. So they keep paying, and the cost creeps into every matter budget.

The Vendor Playbook (And Why It’s Working)

Enterprise AI vendors aren’t moving to token pricing because it’s fair. They’re doing it because their own costs exploded. Training a frontier language model costs $50 million to $200 million. Running inference at scale costs another $10 million per month. Flat seat pricing worked when AI was a feature. Now that it’s the product, the unit economics don’t close.

So vendors are passing the cost to customers, but they’re doing it carefully. The first renewal after launch keeps the old pricing. The second renewal introduces a “hybrid” model with a base fee and a token allowance. The third renewal is pure consumption. By year four, you’re paying 3x what you started with, and the contract language makes it impossible to audit usage or challenge the bill.

The playbook has three steps. First, they bundle the AI feature into your existing subscription so you can’t opt out. Second, they set the initial token allowance high enough that you don’t hit the cap in month one. Third, they bury the overage rate in an appendix and hope you don’t read it until the bill arrives.

It works because most firms don’t have anyone on staff who understands token economics. Your IT director knows servers and licenses. Your CFO knows revenue per partner and overhead ratios. Neither of them knows that a single document review task can consume 50,000 tokens, or that the vendor’s “standard rate” is 4x what you’d pay if you built the same agent on a direct API contract.

We see this every time we run the AI audit for law firms. The firm is paying $8,000 a month for a tool they use twice a week, and the vendor is charging $0.08 per 1,000 tokens when the underlying model costs $0.015. The markup is 5x, and the firm has no leverage because they didn’t negotiate before the renewal.

What Happens If You Wait

If you do nothing, here’s the 18-month timeline. Month one, your current contract renews at the old price. Month six, the vendor emails a “product update” that mentions token-based billing as an option for “high-volume users.” Month twelve, your renewal notice arrives with the new pricing model as the default. Month eighteen, you’re paying triple, and your only option is to eat the cost or rip out the tool mid-year.

The financial hit is obvious, but the operational damage is worse. Your associates stop using the AI because they don’t want to blow the budget. Your intake team goes back to manual triage because the voice agent costs $40 in tokens every time it handles a complex call. Your partners start second-guessing whether to run discovery through the review tool, so they assign it to a junior associate instead. You’re back to the same workflow you had before you bought the software, except now you’re paying $6,000 a month for a tool nobody uses.

The firms that wait also lose negotiating leverage. Once you’re locked into a three-year token-based contract, the vendor has no reason to offer better rates. You can’t threaten to leave because the switching cost is too high. You can’t demand a volume discount because the contract language doesn’t allow renegotiation until year three. You’re stuck.

The firms that act now get a different outcome. They go back to the vendor before the renewal and say, “We’ll commit to another 24 months if you hold the seat-based pricing and give us a fixed token allowance with no overages.” Half the time, the vendor agrees because they’d rather keep the revenue than lose the customer. The other half, the firm walks and builds the same capability on Omni for 60% less.

How Purpose-Built Agents Change the Math

The reason token pricing hurts so much is that general-purpose AI tools aren’t built for legal work. They tokenise everything, even the parts you don’t need. A contract review tool ingests the entire 300-page document, including the signature pages, the exhibits, and the boilerplate. A purpose-built agent ingests only the clauses that matter for your review checklist.

Take the Intake Voice Agent we build on Omni voice. It answers after-hours calls, conflict-checks the caller, and books a consultation. The entire interaction uses 6,000 tokens because the agent is trained on your firm’s intake script, not a general-purpose conversation model. At $0.015 per 1,000 tokens (the rate we pass through from the underlying API), that’s $0.09 per call. Compare that to the $0.80 your current vendor charges for the same interaction, and you’re saving $0.71 every time the phone rings.

The Document Review Agent works the same way. It doesn’t ingest the entire discovery batch. It scans for the clause types you care about (indemnity, limitation of liability, termination rights), extracts those sections, and produces a two-page memo. The token count is 12,000 instead of 150,000. The cost is $0.18 instead of $3. Over 80 documents, you’ve saved $225.60, and the output is faster because the agent isn’t wasting time on irrelevant pages.

The Matter Triage Agent is even cheaper. It reads an intake form submission (2,000 tokens), classifies the practice area (500 tokens), and routes it to the right partner with a one-paragraph brief (1,000 tokens). Total cost: $0.05. Your current CRM charges $0.40 for the same workflow because it’s running a general-purpose model that doesn’t know the difference between a personal injury intake and a contract dispute.

The cost difference compounds. A firm handling 200 intakes, 60 document reviews, and 400 matter triages per month spends $1,840 on a token-metered vendor. The same firm running purpose-built agents on Omni spends $680. That’s $1,160 per month, or $13,920 per year. Over three years, it’s $41,760, and that’s before you account for the time saved by not having associates manually triage every intake.

The Renegotiation You Need to Have This Quarter

If your renewal is more than six months out, you have time. If it’s less than 90 days, you need to move now. Here’s the conversation.

Call your vendor rep and say, “We’re reviewing our AI spend ahead of the renewal. We need to understand the token pricing model and what our projected cost will be under the new structure.” Don’t ask if they’re switching to tokens. Assume they are. Make them explain the rate card, the allowance, and the overage charges.

Then run the math. Take your last three months of usage and multiply by the new token rate. If the number is higher than your current monthly fee, you have a problem. If it’s 50% higher, you have a crisis.

Next, ask for a flat-rate option. Say, “We’ll commit to another 24 months if you give us a fixed monthly price with no token overages.” Some vendors will agree because they’d rather lock in the revenue. Others will refuse because their cost structure doesn’t allow it. If they refuse, you have two options: negotiate a capped token allowance with a reasonable overage rate, or walk.

The capped allowance works like this. You agree to pay $200 per user per month, and the vendor gives you 500,000 tokens per user as part of the base fee. Overages cost $0.03 per 1,000 tokens, and the cap is $400 per user per month. That’s not great, but it’s better than uncapped billing that could hit $800 in a busy month.

If the vendor won’t offer either option, you walk. That’s not a negotiating tactic. It’s a business decision. Paying 3x for unpredictable software is worse than paying nothing and building the capability yourself.

What an Omni Audit Looks Like (And Why It Matters Now)

Most firms don’t know what their AI is costing them until we map it. The Omni Audit takes 60 minutes. We sit with your operations lead, your IT director, and one partner. We walk through every tool you’re paying for, every workflow that’s still manual, and every place where token costs are hiding.

Then we build a cost model. We take your actual usage (intakes per month, document reviews per week, matter triages per day) and calculate what you’d pay under three scenarios: your current vendor’s token pricing, a renegotiated flat rate, and a purpose-built agent on Omni. The difference is usually $80,000 to $250,000 per year for a firm doing $1 million to $25 million in revenue.

We also map the workflow. Where are associates spending four to six hours per week on unbilled admin? Where are after-hours intakes sitting for 12 hours before someone responds? Where are discovery documents piling up because the junior associate is underwater? Those are the places where an agent pays for itself in the first 90 days.

The audit produces three outputs. First, a one-page cost breakdown showing what you’re paying now, what you’ll pay under the new token model, and what you’d pay with Omni. Second, a workflow map showing where agents replace manual work. Third, a 90-day implementation plan with the specific agents we’d build, the integrations we’d connect, and the ROI we’d expect.

We don’t deliver a deck. We don’t schedule a follow-up. We give you the three documents, and you decide whether to move forward. Half the firms book the audit and renegotiate with their vendor using the cost model as leverage. The other half build the agents with us and cut the vendor entirely.

If you’re not sure where to start, we built a checklist that walks through the intake workflow step by step. It’s the same framework we use in the audit, and it’ll show you exactly where token costs are hiding in your current process. You can grab it here: AI Client Intake Checklist for Law Firms. Print it, fill it out, and you’ll know whether your vendor is charging you for work that should cost one-tenth the price.

The Firms That Move First Win

The firms that renegotiate now lock in predictable pricing for another 24 to 36 months. The firms that wait pay triple and lose forecast control. The firms that build their own agents on Omni pay a fraction of either option and own the capability outright.

This isn’t a technology decision. It’s a financial decision. Token pricing is designed to extract maximum revenue from customers who don’t understand the cost structure. The vendors are betting that most firms won’t read the contract, won’t run the math, and won’t have the technical capability to build an alternative.

They’re wrong about the last part. Purpose-built agents are cheaper, faster, and more reliable than general-purpose tools. They don’t punish you for using them. They don’t require a PhD to configure. They don’t lock you into a three-year contract with hidden overage fees.

The window to act is now. If your renewal is in the next six months, book a 60-min Omni Audit and get the cost model before you sign anything. If your renewal already happened, we can still build the agents and cut your token spend by 60% starting next quarter.

The firms that move first will have a two-year cost advantage over their competitors. The firms that wait will spend the next three years paying for software they can’t afford to use. The choice is obvious, but the deadline is real.