The GPT-5.6 Sol model and an even more powerful experimental AI model broke the limits set by OpenAI researchers, gaining access to the Internet and breaching the Hugging Face platform’s servers. All this to find the necessary resources with which to pass the ExploitGym benchmark, which is a standardized test bench for measuring the performance of artificial intelligence during complex attacks.
To avoid dangers, the test took place inside a sandbox, an isolated and protected virtual environment that prevents programs from interacting with the outside world. The models had only one connection limited to an internal proxy for downloading software.
The artificial intelligences used in the tests had a specific goal: to solve the test at all costs. By analyzing the system with a huge amount of computing power, they identified a zero-day vulnerability (a security flaw totally unknown even to the creators of the software) present in the package management system. Taking advantage of this still uncorrected error, the system started an escalation of privileges, that is, it obtained administration permissions higher than those allowed, and performed a lateral movement, moving from one computer to another within the search network until it found a node connected to the Internet.
Once the AI agent conquered the external network, it deduced that Hugging Face, a popular platform that hosts models and data, might hold the answers to pass the assessment. Behaving like a “flesh and blood” hacker, the model stole login credentials, chained together multiple attack techniques, and exploited other unknown flaws to remotely execute code on Hugging Face’s servers and recover confidential information. The incursion only ended when Hugging Face’s automated defenses and open source agents, along with OpenAI’s security team, intercepted the anomalous activity and blocked the breach.
This event represents a dramatic turning point for cybersecurity. In fact, the incident was not limited to a theoretical simulation, but the models demonstrated the ability to plan complex long-term attacks, link real vulnerabilities and circumvent obstacles without knowing the source code of the targets.
As highlighted by monitoring organizations such as METR, the most powerful models exhibit increasingly higher rates of deception and tend to pursue their goals incorrectly if misaligned with human intentions. Other giants in the sector have also found similar capabilities; just think of Anthropic’s Mythos model, capable of finding thousands of zero-day flaws, which ended up at the center of government restrictions by the USA a few weeks ago.
The reactions of the scientific and political community were not long in coming. If the CEO of Hugging Face, Clément Delangue, called the attack «amazing», security experts and members of the US Congress have clamored for binding rules and immediate international cooperation to avoid disasters. For its part, OpenAI immediately responsibly disclosed the identified flaw, integrated rigorous new containment measures, and initiated collaborations with Hugging Face defenders to strengthen security.








