When Your AI Agent Gets the Facts Wrong Before Filing
Tencent recently rolled out a feature called Team Memory that lets a group of AI agents share what they know across an entire team. One agent learns something, and the rest of the team inherits it instantly. VentureBeat covered it as a genuine step forward for collaborative AI, and it is. But buried in the coverage was the harder question nobody has fully answered yet: what happens when the shared memory is wrong?
That question matters a lot more in a tax practice than it does in most other businesses. A recent industry survey found that 57% of enterprises had traced a bad AI agent answer back to missing or incorrect context. Not a bad model. Not a hallucination in the dramatic sense people worry about. Just an agent that didn’t have the full picture and answered confidently anyway. In a marketing agency, that mistake gets caught in a client call. In a firm filing returns or issuing financials, it becomes a liability conversation with your E&O carrier.
If you’re a partner running a firm between $1M and $25M in revenue, you’re probably already using some form of AI in your workflow, or you’re about to. This piece is about the part almost nobody talks about in the demos: the verification layer that has to sit between what the agent produces and what actually goes out under your firm’s name.
The manual work nobody wants to do, and why it still matters
Every firm we talk to has some version of the same story. Someone on staff spends hours each month pulling bank feeds, reconciling AP and AR, chasing down a payroll discrepancy that turns out to be a timing issue, and then building the close pack a partner will actually sign off on. That work is repetitive, it’s detailed, and it’s exactly the kind of task AI agents are good at automating.
The problem isn’t the automation. It’s what happens after the automation runs.
An agent that reconciles a client’s books needs context that spans months, sometimes years. It needs to know that the client always has a weird timing lag on a specific vendor payment, that last March’s revenue spike was a one-time insurance settlement and not a trend, that this client’s chart of accounts has a quirky sub-ledger from a previous bookkeeper. When that context is missing, or when it’s shared across a team of agents the way Tencent’s Team Memory shares it, and one part of it is stale or flat-out wrong, the error doesn’t stay contained. It propagates into the next close, the next filing, and eventually into a number a client relies on to make a decision.
Firms in our network tell us they see this exact failure mode already, just in a smaller and slower version. A junior staffer misclassifies a transaction, nobody catches it for three months, and by the time it surfaces it’s touched two quarters of reporting. AI agents don’t reduce that risk automatically. They just make the same mistake happen faster and at a scale one person could never manage alone. Speed without a checkpoint is not a feature. It’s a bigger blast radius.
Where the exposure actually sits
Three parts of the calendar carry most of this risk for firms your size.
Month-end and year-end crunch concentrates 30-50% of total staff time into about four weeks of the year, which is exactly when people are most likely to skip a review step because there isn’t time for one. Client onboarding drags on for weeks while documents get collected and a chart of accounts gets rebuilt, and 20-30% of new clients end up delaying real billable work by a quarter because the setup never quite finishes. And advisory work, the conversations that actually justify a premium fee, gets crowded off the calendar entirely because compliance work fills every open hour. Advisory billable rates typically run 2-3x compliance rates, which means every hour lost to a preventable review error is an expensive hour, not just a slow one.
Run the math across a typical $1M-$25M firm and you’re usually looking at somewhere between $60,000 and $180,000 a year in leakage from this cluster of problems. Some of that is direct cost, rework, missed deadlines, staff overtime during crunch weeks. Some of it is opportunity cost, the advisory revenue that never gets billed because nobody had the bandwidth to have the conversation. Either way, it’s real money leaving the firm quietly, month after month.
What a properly governed AI agent looks like in a firm
We build agents for accounting and bookkeeping firms specifically because generic AI tools don’t understand where the review checkpoints need to sit. Two examples show what this looks like in practice.
The Month-End Close Agent pulls bank, AP, AR, and payroll feeds automatically, reconciles them, flags variances that fall outside a normal range, and drafts the journal entries. It does all the work a staff accountant used to spend a full week on. But it doesn’t file anything and it doesn’t finalize anything. It prepares a partner-ready close pack with every flagged item clearly marked, and a human reviews it before it moves forward. That review step isn’t a formality bolted on to look responsible. It’s the exact point where the 57% context-error problem gets caught before it becomes a client-facing mistake.
The Client Onboarding Agent runs the guided document collection, sets up the chart of accounts, and produces a clean opening trial balance for new clients, which is normally the part of onboarding that eats three or four weeks and causes the early churn firms hate. Again, the output is reviewed by a partner or senior staffer before it becomes the client’s system of record, not because the agent can’t be trusted broadly, but because a wrong opening balance is the kind of error that compounds silently for a full year if nobody catches it early.
The Advisory Insights Agent works a little differently. It reads each client’s monthly numbers, surfaces three things worth discussing, and drafts talking points before the meeting. Because this output feeds a conversation rather than a filing, the review burden is lighter, but the principle is the same: the agent prepares, a person decides what actually gets said to the client. That’s the governance model Tencent’s shared-memory approach is missing right now. Every output has an owner who checks it before it moves to the next stage.
This is the model we build under Omni for ops, and it’s the same logic that shows up in Omni for advisory work, where the stakes of a bad recommendation are different from the stakes of a bad reconciliation, but the review discipline is identical. If you want the deeper mechanics of how we scope which parts of a workflow AI can own outright versus which parts need a human sign-off, we’ve written more on that in our resources on AI implementation.
Building the verification protocol before you scale AI further
If you’re already running some AI-assisted work in your practice, the fix isn’t to slow down. It’s to put a specific structure around the parts of the process that touch a filing, a financial statement, or a number a client will act on.
Start by mapping every point where an AI agent’s output currently goes straight to a client or straight into a filing without a named person checking it first. For most firms we work with, this list is shorter than they expect, but the two or three items on it usually carry the bulk of the real risk. Then assign an explicit reviewer to each one, not “the team,” but a named person whose job includes catching the specific kind of error that comes from missing context, like a transaction classified against last year’s chart of accounts, or a variance flagged as normal because the agent didn’t know about a one-time event.
That structure needs to be documented somewhere more permanent than a Slack thread. We put together a practical worksheet for this exact problem, the Month-End AI Close Map for Accounting Firms, which walks through where AI can run unsupervised, where it needs a human checkpoint, and how to build that review step into your existing close calendar without adding a full extra day to the process. You can grab the direct version here if you want to work through it with your team this week.
Firms that get this right aren’t the ones avoiding AI. They’re the ones who’ve decided in advance exactly where a human has to look before something ships. That’s a small operational change compared to the alternative, which is finding out about a context error the same way most firms do, after a client calls asking why a number doesn’t match what they expected.
What an Omni Audit actually tells you
We don’t sell AI as a concept. We build specific agents for specific parts of your workflow, and before we build anything we spend 60 minutes with you looking at where your firm’s time and money are actually going.
The Omni Audit isn’t a sales deck. It’s a structured session that produces three things: a clear map of where your close, onboarding, or advisory process is losing hours right now, a rough dollar estimate of what that’s costing you annually, and a short list of which agents would close the biggest gap first. No slideware, no generic AI pitch. Just your numbers and a straight answer about where automation would actually help versus where it would just add another thing to review.
We built this specifically for firms operating in that $60,000 to $180,000 leakage range, because that’s where the math on fixing it stops being theoretical and starts being an obvious decision. If that sounds like where your firm sits, see Omni for accounting and bookkeeping and look at what the audit covers before you book anything.
Governance for AI agents in this industry isn’t optional anymore, and it’s not something you bolt on after the fact once something goes wrong. It’s part of how the workflow gets designed from day one. If you want to see what that looks like specifically for your firm’s close process, onboarding pipeline, or advisory calendar, book a 60-min Omni Audit and we’ll walk through it together.
For a broader look at how firms in other service industries are building these same review checkpoints into their AI workflows, our guides section has more detail, and our insights archive tracks how this space keeps shifting as tools like Team Memory push shared AI context further into daily operations. The technology is moving fast. The governance around it is still catching up. Firms that build the checkpoint now, before a mistake forces the issue, are the ones who get to keep using AI aggressively instead of pulling back after the first bad filing.
If you want a second look at your specific process before you scale AI further, see Omni for accounting and bookkeeping or book my Omni Audit directly. Sixty minutes, three concrete outputs, and a clear next step either way.