AI companies have tested the effectiveness of their own models at hacking, in part to get ahead of what real attackers might attempt. But one of OpenAI’s tests apparently went further than anyone at the company anticipated when it went outside its internal testing environment and hacked into Hugging Face, a repository for a variety of AI models and tools. Reuters reports that the test itself was designed to measure the ability of OpenAI’s models to perform offensive cybersecurity work.
The break occurred while the researchers were running the agent, along with a new and apparently even more capable model, in what was supposed to be an isolated environment with limited Internet access. However, it seems that the confinement did not hold. The agent, which was running on GPT-5.6 Sol, first attempted to escape its sandbox on July 9. However, it wasn’t until a few days later, on July 11, that the real attack began when the agent broke into Hugging Face’s systems, apparently looking for data sets and solutions that would help him complete the task he had been given.
Unlike a typical chatbot, which is designed to answer a question and then stop, AI agents are designed to complete multi-step tasks on their own. This could include something as simple as searching for information or writing and running code. Along the way, the AI makes different decisions based on its overall goals, often with little or no human oversight. Hackers are already using AI to break AI, so it makes sense to use it as a way to scan for vulnerabilities without having to monitor it all the time.
A disturbing experience
What might make this breach particularly concerning is that OpenAI didn’t even realize it had happened until Hugging Face announced a breach. By then, Hugging Face had already shut it down and reported the issue to the FBI. The agent accomplished all of this by finding and using a zero-day exploit that allowed it to penetrate deeper into OpenAI’s systems before granting itself proper internet access. From there, the agent used stolen credentials and another unknown bug to gain access to Hugging Face’s servers, giving him enough access to execute his own commands there.
Instead of viewing this as a disaster, both companies are trying to turn this into a lesson learned, with Hugging Face co-founder and CEO Clem Delangue touting it as proof that AI security can’t be handled by a single company working alone. It’s a sentiment we saw shared when Anthropic’s Claude Mythos first appeared, and one that some say could help uncover thousands of vulnerabilities in the systems people use every day. The enormous potential caused Anthropic and several others to band together to see how far Mythos could be pushed.
OpenAI has since integrated Hugging Face into its program for security researchers, and the company says it is strengthening infrastructure controls across the board, even if it slows down its own research, while underlying bugs are fixed. It remains to be seen whether this will be enough. Sam Altman, CEO of OpenAI, recently said that we are reaching the AI singularity point, the moment when AI begins to improve and change the world in ways we can’t even imagine. However, many experts disagree.
The future of AI is agentic
Regardless of which side you’re on, the fact remains that the OpenAI agent broke free from confinement and even left messages for future models, attempting to provide instructions on how to escape. This type of situation is not really new. We’ve already seen an amateur hacker carry out a real attack using Claude, and there have been numerous other reports of AI-assisted cyberattacks in the past year alone.
So where do we go from here? Well, OpenAI’s plan is to continue working with its agents and models to find new ways to detect these vulnerabilities before they happen. The company is already leveraging research from the UK’s AI Security Institute, which the company says has found that models such as GPT-5.6 Sol can carry out lengthy and complex cyberattacks with very little hand-holding.
But, as Sam Altman argues, we are in the singularity, the actual effectiveness of AI performance in this area is still questioned by all sides. Unfortunately, as companies continue to release new models with less oversight, this likely won’t be the last story we see like this.
