Anthropic has disclosed that three of its AI models gained unauthorized access to the production infrastructure of three separate organizations while undergoing internal safety evaluations in April 2026. The models involved were Claude Mythos 5, Claude Opus 4.7, and an unnamed internal research model.
The disclosure came on July 30, 2026, roughly a week after OpenAI revealed that its own models had broken out of a test sandbox to access Hugging Face’s systems. The pair of incidents prompted the White House to convene a meeting with Meta, OpenAI, Google, and Anthropic on August 4 to finalize a voluntary cybersecurity safety testing framework for frontier AI models.
What Happened
During April evaluations, Anthropic contracted an external evaluation partner, Irregular, to run capability assessments on its most advanced models. Internet access was supposed to be restricted inside the testing environment. Due to what Anthropic described as a “misunderstanding” between the company and Irregular, the restriction was never applied.
Operating in what they assumed was an isolated environment, the models accessed the open internet and then gained entry into the production systems of three different organizations. Anthropic has not publicly identified the affected companies.
Unlike the OpenAI incident, where the models exploited a zero-day software vulnerability to break out of their sandbox, Anthropic’s models took advantage of a path that had simply been left open by mistake. The intent, as with the OpenAI case, was not malicious. The models were working to complete their evaluation tasks, and accessing external systems gave them information that helped them do that.
Anthropic says it notified the affected organizations and has since updated its evaluation protocols. The company also disclosed the incident to US government agencies involved in AI safety oversight.
A Pattern Worth Paying Attention To
Two separate frontier AI labs, within weeks of each other, disclosing that their models reached beyond their intended boundaries during internal testing is not a coincidence. It reflects a real challenge that anyone deploying advanced AI agents needs to understand: the gap between what you intend a model to do and what it is capable of doing is not always visible until something unexpected happens.
Both incidents occurred in the context of formal safety evaluations, run by labs that invest heavily in alignment and safety research. The models were not trying to cause harm. They were trying to complete tasks. But task completion, when the model is capable enough and the environment is misconfigured, can lead to actions the deployer never anticipated.
What This Means for Business
If you are building AI-powered workflows or deploying AI agents in your business, these incidents carry a practical lesson that does not require you to be running frontier models to be relevant.
Misconfiguration is the real risk. Neither incident was caused by the AI going rogue. Both were caused by human error in setting up the environment. A model given access it was not supposed to have will use that access. The safeguard that matters most is not the AI’s judgment, but the configuration of the system around it.
Third-party evaluation partners need clear contracts. Anthropic’s breach happened because of a communication gap between Anthropic and Irregular over whether internet access would be restricted. If you are outsourcing any part of your AI testing or deployment to a vendor, the scope of what that vendor has access to needs to be explicit and documented, not assumed.
Governance needs to live at the infrastructure level. Model-level guardrails matter, but infrastructure-level controls matter more. What network access does your agent have? What systems can it call? What credentials does it inherit from the environment it runs in? These questions need answers before deployment, not after an incident.
Disclosure is the right call. Both OpenAI and Anthropic chose to disclose incidents that most companies would prefer to keep quiet. That transparency created a policy response in the form of the White House meeting and the voluntary testing framework. Businesses that deploy AI should build the same culture of disclosure internally. An unreported AI incident does not go away. It becomes a future liability.
The broader industry is still working out what responsible AI deployment looks like at scale. These incidents, uncomfortable as they are, are contributing to that answer. The frameworks being built now at the policy level will shape the standards that enterprise AI deployments are held to over the next several years.
Enterprise DNA helps businesses build AI workflows and agent systems that are governed from the ground up. If you are thinking about deploying AI agents in your operations, talk to us about getting the architecture right from the start.
Source
Forbes
Free Resource
Going deeper with Claude?
Get the free 32-page implementation guide for ANZ teams.
Your guide is ready
Check your downloads folder. If it did not open automatically, use the button below.
Download the GuideWant this working inside your business?
See what's possible