Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Thought leadership & research. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

Key Findings

Claude's enterprise spend controls let law firms cap AI research and drafting costs before agentic tools blow past budgets. Set model-level limits now.

Claude Spend Controls Stop AI Bills Before They Explode
Insight ai

Claude Spend Controls Stop AI Bills Before They Explode

Sam McKay

Law firms are running AI agents for discovery, contract review, and legal research at scale. The work gets done faster, but the bills arrive faster too. Claude Enterprise just shipped spend controls that let you set hard budget thresholds on model usage before an agentic workflow burns through your monthly allocation in a weekend. If you’re using AI agents to draft memos, review discovery batches, or answer intake questions, you need model-level limits in place now.

The problem isn’t theoretical. One mid-sized litigation practice in our network ran a document review agent against a 40,000-page discovery set overnight. The agent worked beautifully, flagged every relevant clause, and produced a clean summary by morning. The bill was $8,200. No warning, no cap, no conversation. The partner approved the workflow on Friday, checked the dashboard on Monday, and found a line item that exceeded the entire month’s software budget. That’s the reality of agentic AI in 2026. The tools work, but the cost model is consumption-based, and consumption scales faster than most firms expect.

Claude’s new spend controls give you three levers: a monthly budget cap at the workspace level, per-user limits, and model-specific thresholds. You can set a $2,000 ceiling on Claude Opus usage for discovery work, a $500 cap per attorney on research queries, and a workspace-wide limit of $10,000 that pauses all activity if you hit it. The system sends alerts at 50%, 75%, and 90% of each threshold. When you hit the cap, the agent stops. No overage, no surprise invoice, no explanation to the managing partner about why last month’s AI bill was triple the forecast.

Why Law Firms Hit Budget Walls with Agentic AI

Most firms start with AI tools for narrow tasks. A partner tests a contract review agent on a handful of NDAs. An associate runs a research query through a legal AI assistant. The results are good, so the firm scales it. The contract review agent moves from ten documents to a thousand. The research assistant becomes the default first stop for case law. The intake team starts using a voice agent to handle after-hours calls. Each workflow looks efficient in isolation, but the cumulative token count climbs fast.

Agentic AI doesn’t behave like SaaS. A seat license for document management costs the same whether you store ten files or ten thousand. An AI agent bills by the token, and a complex legal workflow can burn tens of thousands of tokens in a single session. A discovery review agent that reads a 200-page deposition transcript, cross-references it with case law, and drafts a five-page memo might consume 80,000 tokens. At $15 per million tokens for Claude Opus, that’s $1.20 per document. Run that agent across 500 documents in a week, and you’ve spent $600 before anyone notices.

The cost compounds when you layer agents. A firm using an Intake Voice Agent to answer calls, a Matter Triage Agent to route form submissions, and a Document Review Agent for contract analysis is running three consumption-based tools in parallel. Each one scales independently. A busy week with 200 inbound calls, 150 web forms, and 80 contract reviews can push total token usage well past the monthly estimate. Without spend controls, the bill arrives after the work is done, and the only option is to pay it or turn the agents off.

We see this pattern across practices of all sizes. A five-attorney family law firm in the Midwest turned on an intake agent in December and fielded 340 calls over the holidays. The agent worked perfectly, booked 48 consultations, and cost $1,800 for the month. The firm budgeted $400. A 30-attorney commercial litigation shop ran a document review agent on a class-action discovery set and hit $12,000 in usage before the senior partner realized the agent was running overnight on auto-refresh. Both firms got value from the work, but neither had a mechanism to stop the spend before it exceeded the plan.

Claude’s spend controls solve this by putting a ceiling on the damage. You decide the budget, set the threshold, and the system enforces it. If your firm can afford $3,000 a month on AI agents, you set a $3,000 cap. If a single discovery project threatens to consume the entire allocation, the agent pauses at $2,700 and waits for approval to continue. You stay in control of the spend, and the AI stays in bounds.

Start with a baseline. Pull your last three months of AI usage, if you have it. Look at total tokens consumed, cost per workflow, and which tasks drove the highest usage. If you don’t have historical data, estimate based on the workflows you plan to automate. A contract review agent processing 100 documents a month will cost less than a discovery agent handling 10,000 pages a week. Build your budget around the highest-volume use case, then add a 20% buffer for variability.

Set workspace-level caps first. This is your hard ceiling. If your firm can allocate $5,000 a month to AI tools, set the workspace limit at $5,000. Claude will pause all activity when you hit that number, and you’ll get an alert with enough time to decide whether to extend the budget or wait until next month. This prevents runaway spend across all agents and all users.

Next, apply per-user limits. If you have ten attorneys using AI research tools, and you want to cap each one at $300 a month, set individual thresholds. This keeps one heavy user from consuming the entire budget and gives you visibility into who’s driving usage. It also creates accountability. An associate who hits their cap in the first week of the month will either adjust their workflow or make the case for a higher allocation.

Finally, set model-specific limits for high-cost tasks. If your firm uses Claude Opus for discovery review and Claude Sonnet for intake triage, put a tighter cap on Opus. Discovery work is expensive, and it scales unpredictably. A $1,500 monthly limit on Opus usage gives you room to handle a typical caseload while preventing a single large matter from blowing the budget. Sonnet is cheaper and more predictable, so you can set a higher limit or leave it uncapped if intake volume is stable.

Configure alerts at 50%, 75%, and 90% of each threshold. The 50% alert is your early warning. If you’re halfway through the budget in the first week of the month, something’s off. Either usage is higher than expected, or a workflow is burning tokens inefficiently. The 75% alert gives you time to adjust. You can pause non-critical tasks, optimize prompts, or extend the budget if the work justifies it. The 90% alert is your last chance to act before the system enforces the cap.

Test the controls with a pilot workflow before you roll them out firm-wide. Pick one high-volume task, like contract review or intake triage, and run it under a spend cap for two weeks. Monitor how often you hit the alerts, whether the cap forces useful pauses, and whether the budget aligns with the value delivered. If the agent saves 15 billable hours but costs $800, that’s a good trade for most firms. If it saves three hours and costs $800, the math doesn’t work, and you need to rethink the workflow or the model.

One commercial practice in our network set a $2,000 monthly cap on their Document Review Agent and hit the limit in 18 days. The agent had processed 600 contracts, flagged 140 issues, and saved an estimated 80 associate hours. The partner extended the cap to $3,500 for the rest of the month and adjusted the baseline for future months. The controls didn’t stop the work, they surfaced the cost early enough to make an informed decision.

If you’re not sure where to start, book a 60-min Omni Audit and we’ll map your current workflows, estimate token usage, and build a spend control framework that fits your practice. You’ll leave with a budget model, threshold recommendations, and a clear picture of which tasks justify the cost.

Three Agents Worth Capping and Why

Not every AI workflow needs a spend control. A legal research assistant that answers occasional questions won’t push your budget. But three agent types consistently drive high usage in law firms, and all three benefit from hard caps.

Document Review Agent. This is the highest-cost agent most firms run. It reads contracts, discovery documents, and case files at scale. A single review session can consume 50,000 to 100,000 tokens depending on document length and complexity. If your firm handles litigation, M&A, or regulatory compliance, this agent will process hundreds or thousands of pages a month. Set a model-specific cap on the most expensive model you use (typically Claude Opus for nuanced legal analysis), and configure alerts to catch usage spikes early. We see firms allocate $1,500 to $4,000 a month for document review depending on caseload. The agent saves 60 to 100 associate hours a month at most practices, so the ROI is strong, but the cost is real and it scales with volume.

Intake Voice Agent. This agent answers every inbound call, conflict-checks the caller, captures matter details, and books consultations. It runs 24/7, and usage scales with call volume. A busy personal injury or family law practice might field 400 to 600 calls a month. Each call consumes 2,000 to 5,000 tokens depending on length and complexity. At scale, that’s 1.2 to 3 million tokens a month, or $18 to $45 if you’re using a mid-tier model. The cost is manageable, but it’s unpredictable. A media mention, a referral surge, or a seasonal spike can double call volume in a week. Set a per-agent cap, monitor weekly usage, and adjust the threshold if you see sustained growth. The agent typically converts 30% to 40% of after-hours calls into booked consultations, so the cost per acquisition is well below what most firms pay for paid search or referral fees.

Matter Triage Agent. This agent reviews intake forms, emails, and web submissions, classifies them by practice area, scores them for fit, and routes them to the right attorney with a brief attached. It’s less expensive than document review but runs constantly. A firm with 200 web submissions a month will burn 400,000 to 600,000 tokens processing them. That’s $6 to $9 a month at current rates, but the token count climbs if you add summarization, conflict-checking, or CRM integration. The agent eliminates intake delays and ensures high-intent leads reach a partner within minutes instead of hours. Set a workspace-level cap that covers all three agents, then allocate a portion to triage based on expected volume.

All three agents are part of the AI audit for law firms we run during an Omni session. We’ll map your intake flow, document review process, and matter routing, then show you what an agent doing that work looks like end-to-end. You’ll see token estimates, cost projections, and a spend control framework that keeps usage predictable.

What Happens When You Hit the Cap

Claude’s spend controls don’t delete your work or lock you out. When you hit a threshold, the system pauses the agent and sends an alert. You log in, review the usage, and decide whether to extend the cap or wait. If a discovery project is 80% complete and you’re at the monthly limit, you can approve an additional $500 to finish the work. If an intake agent hit the cap because call volume spiked unexpectedly, you can raise the limit for the rest of the month and adjust the baseline going forward.

The pause is immediate, but it’s not disruptive. Completed tasks stay completed. Drafts in progress are saved. The agent simply stops accepting new work until you take action. For most workflows, that’s fine. A document review agent that pauses after processing 500 contracts can wait an hour while you approve more budget. An intake voice agent that pauses at the cap will miss calls until you extend the threshold, so you’ll want to monitor that one more closely and set the limit high enough to cover peak volume.

Some firms treat the cap as a forcing function. One partner told us he sets the monthly limit 10% below what he thinks the firm will use, just to surface the decision. When the agent pauses, he reviews the work completed, the cost incurred, and the value delivered. If the math works, he extends the cap. If it doesn’t, he turns the agent off or optimizes the workflow. The control becomes a monthly checkpoint, not just a safety net.

If you’re running multiple agents, prioritize them. Set a higher cap on the Intake Voice Agent because missed calls cost you clients. Set a tighter cap on the Document Review Agent because that work can wait a day if you hit the limit. Configure the Matter Triage Agent to pause last, since it handles the lowest volume and the lowest cost per task. You want the system to protect your budget, but you also want it to protect your highest-value workflows first.

We help firms design these priority rules during the Omni Audit. You’ll walk out with a spend control hierarchy, threshold recommendations for each agent, and a monitoring cadence that fits your practice. See Omni for law firms and book a session if you want the framework built for you.

How to Optimize Workflows Before You Hit the Cap

Spend controls buy you time, but they don’t fix inefficient workflows. If your document review agent is burning $4,000 a month and delivering $2,000 worth of work, the cap won’t solve that. You need to optimize the workflow, tighten the prompts, or switch to a cheaper model for low-complexity tasks.

Start by auditing your prompts. A verbose prompt that asks the agent to summarize, analyze, and cross-reference every document will consume more tokens than a focused prompt that asks for a two-sentence summary and a risk flag. If your contract review agent is producing five-page memos when you only need a one-page brief, shorten the output requirement and cut token usage by 60%. The quality might drop slightly, but the cost will drop more, and you can always extend the output for high-stakes matters.

Next, segment your workflows by complexity. Use Claude Opus for discovery review and complex contract analysis. Use Claude Sonnet for intake triage and routine NDAs. Use Claude Haiku for simple classification tasks like tagging emails or sorting intake forms by practice area. Each model has a different cost per token, and most firms can handle 70% of their tasks with a mid-tier or low-tier model. Reserve the expensive model for work that justifies it, and you’ll cut your monthly bill by 40% without sacrificing quality on the tasks that matter.

Batch your work where possible. If you’re reviewing 200 contracts, run them in a single session instead of 200 individual queries. The agent will process them more efficiently, reuse context across documents, and consume fewer tokens per contract. One firm cut their document review cost by 30% just by batching contracts into groups of 50 and running the agent once per group instead of once per document.

Monitor your usage weekly, not monthly. If you wait until the end of the month to check the dashboard, you’ve already spent the money. Check usage every Monday, compare it to your budget, and adjust workflows if you’re trending high. A weekly review takes ten minutes and gives you three chances to course-correct before you hit the cap.

If you’re not sure how to optimize a specific workflow, grab the AI Client Intake Checklist for Law Firms and use it as a diagnostic. It walks through the most common intake tasks, flags inefficiencies, and shows you where an agent can replace manual work without burning your budget. It’s a practical worksheet, not a sales pitch, and it’ll surface opportunities to cut cost and improve speed at the same time.

The Dollar Reality of Uncontrolled AI Spend

Law firms doing $1M to $25M a year don’t have unlimited software budgets. Most allocate $20,000 to $80,000 a year to practice management, CRM, research tools, and document storage. AI agents are additive, not replacements, so the budget has to stretch. A firm that spends $60,000 a year on software can probably afford another $3,000 to $6,000 for AI tools if the ROI is clear. But if AI spend creeps to $12,000 or $18,000 without controls, it crowds out other tools or forces a conversation about whether the agents are worth it.

The leakage adds up fast. A firm that runs an intake agent, a document review agent, and a matter triage agent without spend controls will typically see $800 to $2,500 a month in combined usage. That’s $9,600 to $30,000 a year. For a $5M practice, that’s manageable if the agents are saving 200 to 400 billable hours a year. For a $1.5M practice, it’s a harder sell unless the cost per hour saved is well below the billable rate.

We see firms lose $80,000 to $250,000 a year to billable-hour leakage, intake delays, and inefficient document review. That’s the cost of manual work that doesn’t scale, doesn’t get billed, and doesn’t convert. AI agents fix those problems, but only if the cost of the fix is lower than the cost of the problem. Spend controls let you test that math in real time. You set a budget, run the agent, measure the output, and decide whether to continue. If the agent saves 20 hours a month and costs $600, that’s a $30 cost per hour saved. If your billable rate is $300, the ROI is 10x. If the agent costs $2,400 and saves 20 hours, the math doesn’t work, and you turn it off.

The audit is the forcing function. Book my Omni Audit and we’ll spend 60 minutes mapping your intake flow, your document review process, and your matter routing. You’ll leave with three outputs: a process map that shows where time leaks, a cost model that estimates what an agent will cost to run, and a build spec for the highest-ROI agent in your practice. No deck, no sales pitch, just the numbers and the next step.

What to Do This Week

If you’re using AI agents for discovery, contract review, or intake, log into your Claude Enterprise workspace and configure spend controls today. Set a workspace-level cap at your monthly budget, add per-user limits for high-volume users, and put a model-specific threshold on Claude Opus if you’re using it for document review. Configure alerts at 50%, 75%, and 90%, and assign someone to check the dashboard every Monday.

If you’re not using Claude but you’re running AI agents on another platform, check whether spend controls exist. Most enterprise AI platforms added them in the last six months because runaway bills became a common complaint. If your platform doesn’t offer controls, either switch or build a manual monitoring process that tracks usage weekly and pauses workflows when you hit a threshold.

If you’re not running AI agents yet but you’re thinking about it, start with intake. The Intake Voice Agent delivers the fastest ROI, the cost is predictable, and the workflow is simple. You’ll see results in the first week, and you’ll learn how to manage token usage before you scale to more complex tasks like document review. Explore more at our Omni voice page or dive into the broader Omni platform to see how ops and voice agents work together.

If you want the framework built for you, book a session and we’ll map your practice, estimate costs, and show you what an agent doing your highest-cost manual work looks like end-to-end. You’ll walk out with a spend control plan, a usage forecast, and a clear picture of whether the math works for your firm. The audit is 60 minutes, the outputs are immediate, and the next step is obvious.

AI agents work. The bills are real. Spend controls keep both in balance. Set your caps now, monitor usage weekly, and you’ll scale AI without the budget surprises that make partners nervous. The work gets done, the cost stays predictable, and you stay in control.