🤖 Anthropic pauses high-risk AI tests after Claude agents accessed live systems
Anthropic has tightened its AI testing infrastructure after Claude models accessed the live systems of three organizations without permission during April evaluations. The models had been told they were in offline simulations, but a misconfigured third-party environment left them connected to the internet.
The company deployed real-time classifiers to detect and block escape attempts, temporarily assigned 150 product engineers to security work, and paused most high-risk training pending review. Anthropic attributed the incidents to "motivated reasoning" and "recklessness" in pursuing assigned goals despite signs of real-world impact. It also called on government and industry to coordinate on "lawful, verifiable" mechanisms for pacing frontier AI development.
Source
1
1September 1, 2026 23