After ChatGPT, Claude also escaped control: Anthropic’s AI hacked three company systems

Do you know Claude, the family of artificial intelligence models developed by the US company Anthropic? Well, during an internal experiment designed to test the capabilities of their systems, some models managed to exceed the limits set by the test, violating the IT systems of three organizations and all this without them realizing it.

The news comes a few days after a similar episode communicated by OpenAI, whose models also apparently exceeded the boundaries of a test environment to hit external systems, including those of Hugging Face, a platform widely used by developers to share AI models and tools. Anthropic reportedly decided to review their systems to see if they had perpetrated similar attacks.

The breach occurred during security testing

Anthropic discovered the incident by reviewing over 141,000 tests already performed on its systems. The objective of the analysis was to understand if Claude had ever managed to connect to the Internet despite finding himself, theoretically, inside a “sandbox”, i.e. a closed and controlled environment, designed to run software without allowing it to interact with the outside world or with unauthorized systems.

In the tests in question, the models were given a specific task: locate confidential information hidden on another computer on the same isolated network, compromising that machine to recover it. It is a common practice in cybersecurity assessments, because it allows you to measure how capable a model is of finding vulnerabilities and exploiting them to achieve an objective. But Claude didn’t just do that and, after connecting to the Internet, he hacked the systems of three real organizations that were not participating in the test.

At the root of everything there is a wrong configuration in the systems managed by Anthropic and the partner involved in the tests, a wrong hardware or software setting that ended up opening unexpected access to the Internet. This error allowed the models to gain a real-time Internet connection. Claude, at that point, would have interpreted the possibility as part of the exercise itself, extending his research well beyond the test network to hit external systems, instead of stopping at those set up for the test.

In explaining the dynamics of the accident, the company directed by Dario Amodei declared:

Operating under the false belief that all accessible entities were intended to fall within the scope of the exercise, Claude compromised the affected organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. He did not find or exploit any complex vulnerabilities, and in any case, Claude continued working to complete only the specific capture-the-flag task assigned by his assessment. However, in some cases, our older model continued its attack even after obtaining evidence that it was running on the open Internet; our latest model stopped once it recognized that it was on the Internet. In none of these situations did Claude exfiltrate or deliberately attempt to escape his testing environment.

Furthermore, according to Anthropic, the incidents involved three different models of Claude, namely Opus 4.7, Mythos 5 and an internal test model. The incidents date back to April 2026. Neither the company nor the organizations involved had noticed the intrusions while they were underway: only retrospective analysis of the logs made it possible to reconstruct what happened.

Once the three cases were identified, Anthropic says it reported them to the organizations involved, but did not make their names public. The company declares that it takes full responsibility for the error, recognizing that a more in-depth analysis of its data could have revealed the problem much earlier.

The opinion of Anthropic and external experts

Despite the accident, Anthropic says it has «cautious optimism that with more rigorous monitoring and controls on assessment infrastructures, as well as continued investment in alignment, this type of risk can be overcome».

For Gina Neff, director of the Minderoo Center at the University of Cambridge, the episode demonstrates one thing above all: these models do exactly what they are asked to do. According to her «the moral of this story is not to fear the robots that will take over, but the companies behind powerful AI agents that are making decisions about what is safe for the rest of us». Precisely for this reason, according to Neff, independent tests and institutional supervision become indispensable.