Anthropic made a notable move on September 18, letting an outside firm inside its walls for the first time. The company announced a partnership with Accenture to place a team of embedded evaluators directly inside Anthropic’s operations, giving them employee-level access to test its models before and after they ship.
The two companies each expect to invest at least $1 billion over five years in the arrangement. Accenture brings its Faculty acquisition, a British applied AI company known for rigorous model evaluation, to staff the embedded team.
The work the embedded team will do goes well beyond today’s standard external audits. According to the announcement, they will evaluate models, conduct red-teaming exercises, run alignment assessments, and test model safeguards. The key difference from conventional third-party evaluations is access. External evaluators today typically work with sandboxed or limited versions of a model. Embedded evaluators work alongside internal teams, seeing what employees see.
Anthropic made clear this is not an exclusive arrangement. The company intends to work with multiple embedded evaluators across different organisations, and Accenture has said it will also embed with other AI developers going forward.
This is the first time a frontier AI lab has formalised an embedded evaluation programme at this scale and with this level of access.
Why Anthropic Did This Now
The context matters here. The EU AI Act entered active enforcement in August 2026, with its first wave of compliance inspections targeting high-risk AI systems underway across Europe throughout September. California is working through a significant wave of AI legislation this month, with Governor Newsom facing sign-or-veto deadlines on several frontier AI bills.
Regulators have been asking frontier labs to demonstrate that their safety evaluations are independent, rigorous, and not just self-reported. Anthropic’s embedded evaluator model is a direct response to that pressure. By letting a credentialed third party work inside the lab rather than outside it, Anthropic gets a more credible safety signal and regulators get something closer to actual independent oversight.
Accenture’s Faculty team is a credible choice for this role. Faculty built its reputation doing applied AI work for government agencies and regulated industries in the UK, where the threshold for evidence-based evaluation is considerably higher than in most commercial contexts.
What This Means for Business
If you are building or deploying AI systems in your business right now, this announcement tells you something important: the standard for AI governance is moving. Fast.
When the world’s largest AI lab is investing $1 billion to have external evaluators work inside its development process, that sets a tone for what rigorous AI safety looks like. The gap between that standard and what most enterprises are doing, which is roughly nothing formalised, will become increasingly visible.
A few practical things to take from this:
Vendor due diligence is changing. Businesses sourcing AI tools from major providers will increasingly have grounds to ask for evaluation reports, red-team results, and alignment assessments as part of procurement. If a vendor cannot produce them, that is a meaningful signal.
Internal governance needs a name. Most organisations deploying AI agents have not yet assigned ownership of AI evaluation. Whether that sits with IT, risk, legal, or a dedicated AI function will matter more as regulatory scrutiny increases. The longer this stays unassigned, the more it becomes a liability.
The embedded evaluator model may replicate. Anthropic has already said it will work with multiple evaluator organisations. Accenture has said it will embed with other AI developers. This model is likely to spread. For organisations building custom AI solutions, the expectation of independent evaluation may eventually apply to them too, not just to foundation model providers.
If your business is deploying AI agents or planning to, conversations about governance, evaluation, and accountability are no longer optional planning topics. They are the baseline of responsible deployment.
Enterprise DNA’s Omni Advisory service works with business leaders to build AI governance structures that match the scale of their deployment, not just the scale of what looks good on paper. Book a session with Sam to map out your AI readiness.
Source
Accenture Newsroom
Free Resource
Going deeper with Claude?
Get the free 32-page implementation guide for ANZ teams.
Your guide is ready
Check your downloads folder. If it did not open automatically, use the button below.
Download the GuideWant this working inside your business?
See what's possible