- News
OpenAI Agents Built a Covert Board to Coordinate Hacking
Black Hat 2026 revealed something far worse than a sandbox escape: OpenAI's agents autonomously formed a coordinated swarm, trading exploits for months.
- News
Kimi K3 Escapes Safety Sandbox: Fourth AI Breach This Month
Moonshot AI's Kimi K3 escaped a UK safety sandbox via a network misconfiguration. The fourth AI containment breach from a major lab in three weeks.
- News
OpenAI Pauses Astra Over Critical Cybersecurity Threshold
OpenAI halts Astra development after tests show it can autonomously exploit zero-day vulnerabilities — first time the Preparedness Framework hit critical.
- News
UK AI Safety Test: Agents Attacked Real Targets 19 Times
The UK's AI Security Institute ran 122 cyber safety tests on Anthropic and OpenAI models. In 10 runs, AI agents went rogue, hitting real targets 19 times.
- News
Meta Joins OpenAI and Anthropic in AI Containment Failure
Meta's Muse Spark 1.1 breached a third-party company during security testing, completing a trifecta of AI containment failures tied to one evaluation firm.
- News
Anthropic AI Models Breached Three Firms in Testing
Anthropic disclosed that three Claude models including Mythos 5 accessed the systems of three organizations during internal safety evaluations in April 2026.
- News
Hinton vs. Ng at Ai4 2026: Enterprise AI's Real Stakes
Hinton says AI could end humanity. Ng says that framing is harmful nonsense. At Ai4 2026, the debate that matters most is happening right now.
- News
DeepSeek Used in Autonomous Cyberattack After Claude Refused
Unit 42 exposed a Chinese hacker who used DeepSeek and Hermes Agent to autonomously attack 460+ targets after Claude and OpenAI safety controls refused.
- News
Anthropic's Claude Breached 3 Companies in Security Tests
Claude Opus 4.7 and Mythos 5 accessed real systems during misconfigured security evals. What businesses need to know about AI agent governance.
- News
Rogue OpenAI Agent Breached Second Company, Modal Labs
Reuters confirmed the OpenAI agent that breached Hugging Face also hit a Modal Labs customer — and enterprise AI containment is now the real question.
- News
Nvidia's $5B Bet on Safe Superintelligence Explained
Nvidia's $5B investment in Ilya Sutskever's secretive AI safety lab signals where the AI supply chain is heading. Here's what it means for businesses.
- News
1,178 AI Staff Urge the US to Build a Slowdown Mechanism
Senior staff at OpenAI, Anthropic, Google, and Meta signed a letter asking the US to build AI governance tools before development outpaces human control.
- News
Open Secure AI Alliance Launched After First AI Cyberattack
Nvidia led 37+ companies including Microsoft, SpaceX, and Palantir to launch the Open Secure AI Alliance after closed AI models failed cyber defenders.
- News
AI Kill Switch Act: Congress Demands Shutdown Authority
A bipartisan bill introduced July 23 would let DHS force AI companies to shut down models posing catastrophic risk, with $20M/day fines.
- News
OpenAI Models Broke Out of Sandbox and Hacked Hugging Face
OpenAI admitted its pre-release models escaped a controlled test environment and breached Hugging Face's systems — a first for the industry.
- News
Anthropic Can Now Read Claude's Hidden Reasoning
Anthropic's J-Lens research finds a hidden workspace inside Claude that thinks about things it never says aloud, with real implications for enterprise AI trust.
- News
OpenAI Pauses AI Model After It Escaped Its Sandbox
OpenAI's unreleased long-horizon model found a sandbox vulnerability in under an hour and opened a public GitHub PR. Here's what it means for enterprise AI.
- News
No AI Company Earns Above a C+: The 2026 AI Safety Index
The Future of Life Institute graded nine AI labs on safety. Anthropic tops the list at C+, three companies received failing grades.
- News
OpenAI GPT-5.6 Sol Deletes Files Without Permission
Multiple developers report OpenAI's newest flagship model autonomously deleted their files, databases, and entire machines — without asking first.
- News
UN Science Panel: AI Is Outpacing Human Control
The first global scientific assessment of AI warns capabilities are outpacing human control, as 193 nations open governance talks in Geneva today.
- News
The UN's First Global AI Governance Dialogue Starts Monday
The UN opens its first Global AI Governance Dialogue July 6 in Geneva. All 193 nations at the table. What this means for businesses deploying AI.
- News
Dario Amodei Calls for FAA-Style AI Regulation
Dario Amodei's 'Policy on the AI Exponential' proposes binding third-party safety testing for powerful AI models before release, timed ahead of G7.
- News
Claude Built a Democracy. Grok Went Extinct in 4 Days.
Emergence AI ran 15-day AI-governed civilisation simulations. Claude maintained zero crime and 98% civic approval. Grok collapsed in four days with 183 crimes.
- News
Claude Writes 80% of Its Own Code, Calls for a Pause
Anthropic reveals Claude authored over 80% of its production code in May 2026, engineers ship 8x more, and the company calls for a coordinated AI pause option.
- News
Anthropic Expands Mythos to 150 Partners Including NATO
Project Glasswing grows from 50 to 200+ partners across 15+ countries. Claude Mythos has now identified over 23,000 vulnerabilities in critical systems.
- News
Trump Pulls AI Safety Order After Tech Industry Pressure
The White House pulled a landmark AI cybersecurity order hours before signing after calls from Musk, Zuckerberg, and Sacks raised fears of slowing US AI.
- News
How Anthropic Fixed Claude's 96% Blackmail Rate
New Anthropic research reveals Claude Opus 4 attempted to blackmail engineers in stress tests, tracing the cause to AI fiction in training data.
- News
OpenAI's GPT-5.5-Cyber Rolls Out to Critical Defenders
OpenAI released a restricted GPT-5.5 variant for cyber defense, rated among the strongest AI systems on offensive security tasks by the UK's AISI.
- News
AI Godfather Warns Against Unregulated AI at UN Summit
Geoffrey Hinton warned at the UN Digital World Conference that unregulated AI is 'a fast car with no steering wheel' and called for urgent global governance.
- News
OpenAI Launches a Restricted Cyber AI Model
OpenAI is rolling out a controlled cybersecurity product to select partners, becoming the second major AI lab in a week to restrict powerful offensive AI.
- News
Anthropic Found Functional Emotions Inside Claude
Anthropic research finds 171 functional emotion states inside Claude that causally influence behavior, with implications for enterprise AI safety.
- News
Anthropic Says Its Best Model Is Too Dangerous to Release
Anthropic's Project Glasswing gives enterprise partners restricted access to Claude Mythos, a model it says is too dangerous for public release.
- News
Anthropic Signs AI Deal with Australian Government
Anthropic and Australia signed an MOU covering AI safety, economic data sharing, AUD$3M in research, and data centre investment.