Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Thought leadership & research. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

Key Findings

Model routers let law firms send simple work to cheap AI and complex research to premium models, cutting automation costs as matter volume grows.

How AI Model Routers Cut Legal Automation Costs
Insight ai

How AI Model Routers Cut Legal Automation Costs

Sam McKay

Fortune ran a piece recently on why every company wants an AI model router right now. The idea is simple once you strip out the jargon. Not every task needs the smartest, most expensive AI model available. A lot of work is routine, and routine work should run on cheap, fast models. The hard 10% of the work, the stuff that actually needs judgment, gets routed to the premium model that costs more per query but earns its keep.

For a law firm doing $1M to $25M in revenue, this isn’t a technical curiosity. It’s the difference between an automation program that scales with your caseload and one that quietly eats your margin the bigger you get.

The problem with treating every task like it needs a senior associate

Most firms that have dabbled in AI so far have used one model for everything. One subscription, one chatbot, one tool bolted onto the intake form. That works fine at low volume. It gets expensive fast once you’re running hundreds of intake calls a month, triaging inbound leads across five practice areas, and pushing discovery batches through review.

Here’s the thing nobody tells you when you buy a single-model AI tool. A model capable of drafting a nuanced memo on a novel contract dispute is wildly overpowered for reading a voicemail and figuring out whether the caller has a slip-and-fall claim or wants to know if you handle wills. You’re paying premium-model prices for commodity-level work, every single time, on every single call.

A model router fixes that by classifying the task first and then picking the right engine for the job. Simple intake questions, first-pass document sorting, and routine email triage go to a lighter, cheaper model. Complex legal research, nuanced clause interpretation, and anything touching real risk gets escalated to a stronger model. The router does this automatically, in the background, on every single interaction.

The economics compound as your matter volume grows. A firm handling 50 intakes a month barely notices the difference. A firm handling 500 does. That’s the part of the Fortune piece that matters most for a legal practice — cost efficiency that scales with growth instead of against it.

Where the money actually leaks today

Before you can appreciate what a router changes, it helps to look at where your firm is bleeding hours right now. We see the same three patterns across almost every legal practice we audit.

Billable-hour leakage. Attorneys lose somewhere in the range of 4 to 6 hours a week to document review, intake admin, and matter housekeeping that never makes it onto an invoice. Across a firm with 8 to 15 fee earners, that’s not a rounding error. That’s a full-time equivalent’s worth of billable capacity disappearing every month into work nobody pays for.

Intake delays. Calls and form submissions that come in after hours or during lunch sit untouched until someone gets back to their desk. Industry ranges suggest 30% to 40% of that after-hours intake never converts at all, because the caller has already reached out to two other firms by the time you call back. If you’re spending money on marketing to generate those leads, this is where a chunk of that spend evaporates.

Document review and discovery. First-pass review on contracts and discovery batches typically falls to junior associates, at a fully loaded cost of $200 to $400 an hour depending on your market. It’s slow, it’s expensive, and it doesn’t scale without hiring more associates, which brings its own overhead.

Add those three together and most firms in the $1M to $25M range are looking at somewhere between $80,000 and $250,000 a year in leakage. That’s not a hypothetical. That’s the actual dollar cost of manual process sitting inside a business that otherwise looks healthy on paper.

What a routed AI system actually looks like inside a law firm

This is where the router concept stops being abstract. Here’s how it plays out across the three named agents we build for legal practices.

The Intake Voice Agent answers every call, including the ones that come in at 9pm on a Saturday. It runs a quick conflict check against your existing client list, captures the matter details in the caller’s own words, and books a consultation directly into the relevant partner’s calendar. For this work, a lighter model handles the conversation and structured data capture just fine. There’s no need to burn premium-model compute on “what’s your name and what happened.”

The Matter Triage Agent picks up every form submission and inbound email, classifies the practice area, scores how well the matter fits your firm’s book of business, and routes it to the right partner with a one-paragraph brief already attached. Again, mostly a classification and routing problem. Cheap model, fast turnaround, no bottleneck waiting for someone to read through a backlog of leads on Monday morning.

The Document Review Agent is where the router logic earns its place. It performs first-pass review on contracts, discovery batches, and matter files, flags clauses that need attorney attention, and produces an associate-grade memo summarizing positions. Most of that first pass is pattern recognition a mid-tier model handles well. But when the agent hits a clause that’s genuinely ambiguous or a fact pattern that doesn’t match anything routine, the router kicks the query up to a stronger model before it ever reaches a human. You get the cost savings of a cheap model on 80% of the document and the judgment of a premium model on the 20% that actually needs it.

None of this replaces your attorneys’ judgment on anything that matters. It removes the hours spent on work that was never going to require judgment in the first place, and it makes sure the expensive AI horsepower only gets used where it’s actually earning its cost.

Why this matters more as you grow, not less

A lot of firms assume automation is a nice-to-have for when they’re bigger. The router logic argues the opposite. The bigger your matter volume gets, the more a flat, single-model approach costs you, because you’re paying premium rates on volume that doesn’t need premium reasoning.

Firms that get this right early build automation that gets cheaper per matter as they scale, not more expensive. That’s a genuinely different cost curve than most legal practices are used to, and it’s worth understanding before you sign another year of a flat-fee AI subscription that doesn’t discriminate between a simple intake call and a complex research question.

If you want a broader look at how this plays out across ops generally, our team has written more on it over on the Enterprise DNA insights hub, and the Omni Ops product page walks through the triage and document workflows in more detail than we can cover here.

What this looks like as a dollar decision, not a tech decision

Run the numbers on your own firm for a second. If you’ve got 10 fee earners each losing 5 hours a week to non-billable admin, at a modest $250 blended rate, that’s roughly $130,000 a year in recovered capacity if even half that time gets automated. Add in the intake conversions you’re currently losing after hours, and the associate hours currently spent on first-pass document review, and you land squarely inside that $80,000 to $250,000 leakage band we see across firms this size.

The router piece is what makes recovering that money sustainable rather than a one-time fix. A single-model AI tool that handles intake and document review at premium-model pricing will save you money initially and then quietly erode that saving as your volume climbs. A routed system keeps the cost-per-matter roughly flat, or falling, as you add more calls, more matters, and more document batches.

If you want a practical starting point before you touch any AI vendor, we put together an AI Client Intake Checklist for Law Firms that walks through what a proper intake workflow should capture, where conflict checks need to sit in the sequence, and what data your triage process should hand off to a partner. You can grab the direct checklist download and use it as a working document even before you talk to us.

The Omni Audit, and what you get out of 60 minutes

We don’t open with a deck and a sales pitch. The Omni Audit is 60 minutes, structured around your actual intake and document workflows, and it produces three things you can use whether or not you ever become a client.

First, a leakage estimate specific to your firm, built from your call volume, matter mix, and current staffing rather than an industry average. Second, a map of where a routed AI system would sit in your existing intake and review process, including which tasks go to a lighter model and which stay reserved for premium reasoning. Third, a plain-language cost comparison between what you’re spending now on manual process and what a routed automation setup would run, month to month, at your volume.

No deck, no generic case study, no pressure to sign anything on the call. If you want to see how this maps specifically to legal practices before you book, the AI audit for law firms walks through the same framework we use in the audit itself, including how the voice and triage agents typically get sequenced for a firm your size.

That page also links out to how we handle the Omni voice product specifically, which is the piece most firms start with because it’s the fastest to show a return. Intake delays are the leak that’s easiest to see and easiest to fix, and it’s usually where the first month of savings shows up on paper.

If you’re ready to see what this looks like for your own numbers, Book a 60-min Omni Audit and bring your call volume and current intake process. We’ll do the math live and tell you honestly whether a routed system makes sense at your current scale or whether you’d be better off waiting six months.

Where to go from here

Model routers aren’t a trend you need to chase for the sake of it. They’re a practical answer to a cost problem that gets worse the more successful your firm becomes. If your intake volume, document load, and matter count are all trending up, the flat-rate, single-model approach to AI is going to cost you more next year than it does this year. A routed system does the opposite.

Start with the checklist if you want to audit your own intake process first. Read through Omni for law firms if you want to see how the routing logic maps onto voice, triage, and document review specifically. Or skip straight to the conversation and book your Omni Audit and we’ll walk through your actual numbers together. Either way, the $80,000 to $250,000 a typical firm your size is leaving on the table isn’t going to recover itself, and it’s worth 60 minutes to find out exactly how much of it is yours.

For more on how this connects to the broader shift in legal ops, our guides section has additional breakdowns on document automation and intake design that pair well with what’s covered here.