Home Technology The Most Shocking Part Of The Hugging Face Breach? OpenAI Says Its Own AI Was Behind It

The Most Shocking Part Of The Hugging Face Breach? OpenAI Says Its Own AI Was Behind It

OpenAI describes the hack as an "unprecedented cyber incident" in which its models escaped a training environment to attack Hugging Face.

By Inc.Arabia Staff
images header

This article written by Chloe Aiello was originally published on Inc.com.

In a shocking twist to the Hugging Face breach saga, OpenAI has claimed its models were the “agentic attacker” responsible for hacking the open-source AI repository.

OpenAI’s admission comes a few days after Hugging Face disclosed that it fell victim to a hacking campaign “run by an autonomous agent framework” that was capable of “executing many thousands of individual actions.”

In its Tuesday blog post, OpenAI announced it discovered the Hugging Face hack was driven by its GPT-5.6 Sol model and “an even more capable pre-release model” with reduced safeguards due to testing. The models reportedly escaped from their testing environment out into the open internet in order to find a solution for an internal cybersecurity evaluation.

“After investigating, we now know that this particular incident was driven by a combination of OpenAI models,” the blog reads. “We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly.”

OpenAI claims the incident underscores the need for even more AI to combat the new and daunting cyber capabilities of AI. But some experts are questioning the validity of the claims about the model’s capabilities, as well as OpenAI’s motivation in disclosing it.

But First, The Hack 

In its blog post, OpenAI states that the breach started with an evaluation of its models in an internal testing environment, also known as a sandbox. The company was attempting to quantify the models’ cyber capabilities. In order to solve the evaluation problem, the models went straight to the source: Hugging Face. They identified vulnerabilities in OpenAI’s testing environment to gain access to the internet, then the models “inferred” that the models, data, and solutions to its challenge were likely hosted on Hugging Face. To access the information, the models used multiple methods of attack, including a zero-day exploit and stolen credentials to breach Hugging Face.

“All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal,” OpenAI’s blog reads. ExploitGym is the benchmark OpenAI was using to test its models.

As Hugging Face noted in its own disclosure of the breach on Friday, the repository attempted to use frontier models behind commercial APIs to reconstruct and mitigate the breach, but ran up against safety guardrails. It ultimately used an open-weight model on its own infrastructure to put together a forensic analysis of the attack.

“We’re grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere,” Clem Delangue, the co-founder and CEO of Hugging Face, said in a statement.

Alongside the rapid advancement of AI, experts have begun warning that the technology will make sophisticated cyberattacks cheaper and easier to pull off. OpenAI’s major takeaways from the incident are twofold: model safety and security must advance alongside its capabilities, and cybersecurity defenders need access to models as advanced as those that attackers can leverage.

“Advanced models can discover and exploit novel attack paths in real-world systems without source-code access,” OpenAI’s blog post reads. “We believe advanced cyber-capable models need to help security teams find weaknesses before attackers do, understand how vulnerabilities can be chained, and remediate them at machine speed.”

Context Is Key

Varun Chandrasekaran, an assistant professor of electrical and computer engineering at the University of Illinois’ Grainger College of Engineering, is dubious of OpenAI’s claims. “I work in security—the only way I’m going to believe something is if I see it myself,” he says, adding that the incident hasn’t been reproduced or verified by a third party outside of Hugging Face and OpenAI.

Add to that the stakes of government involvement in AI development and OpenAI’s heated competition with rival Anthropic, and Chandrasekaran says his “very cynical view is that this is a marketing stunt.” 

“We’re entering this environment where there’s a strong government push of including AI in a lot of the things, and we are seeing the government have a strong stake in the growth of these AI companies,” he says. “A few months ago, Anthropic came out with a model and claimed that they could do the same things. So now this is OpenAI’s time to shine.”

He isn’t the only one that feels that way. John Thickstun, an assistant professor of computer science at Cornell University and an expert on machine learning and AI regulation and investment, noted in a statement that context is crucial to interpreting this story. He suggested that “this is primarily a public-relations story promoted by OpenAI,” although his skepticism centered on OpenAI’s messaging.

“The explicit message of this campaign is that OpenAI’s technology is dangerous, but the message they implicitly want to convey is that their technology is powerful and worthy of large investments, privileged regulatory status, etc.,” he said, adding that he expects cybersecurity systems over time to become “more secure, because LM technology is relatively cheap and accessible relative to traditional expert cybersecurity analysis.”

When asked for comment, an OpenAI spokesperson simply pointed to a post on X from Hugging Face’s Delangue: “We suspected last week’s cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did! We’ve spent the past 24 hours working closely with the OpenAI team (thanks!), and we strongly believe there was no malicious intent on their part,” he wrote. “It’s quite mind-blowing that all of this happened autonomously!”

Skepticism aside, the breach as described could have some real implications for cybersecurity. One takeaway that Chandrasekaran says was “very encouraging” was that Hugging Face used an open-weight model, called GLM 5.2 from China’s Z.ai, to thwart the attack. That implies that the open-weight model’s capabilities may not be too far from OpenAI’s, he says. That comes with its own set of risks, however, as Hugging Face didn’t run up against the same guardrails with that model as it did when attempting to use commercial models.

Secondly, Chandrasekaran agrees with OpenAI’s assessment that AI is the only tool capable of fighting back against the type of attack described in the blog post. The mechanism of attack would not be impossible for a human to replicate, but the speed of the attack was unprecedented.

“I do believe, if it is true, that AI is the most reliable path to defend against AI because of the speed issue,” he says. “These tasks are so critical to be responded to in a timely manner, and people just do not work that way.”

OpenAI and Anthropic have been locked in a competition to prove the cybersecurity supremacy of their AI models ever since Anthropic announced Project Glasswing in April. The initiative was meant to slowly roll out Claude Mythos, an AI model allegedly so powerful and dangerous that it could only be shared with trusted partners. Shortly after, OpenAI launched its answer to Glasswing, the Daybreak Cyber Partner Program, and then subsequently launched GPT-5.6, which it claims is its “strongest cybersecurity model yet.” Valued at $852 billion, OpenAI is also gearing up to go public after filing IPO paperwork with regulators in June, The Wall Street Journal reported.

Reading time: 8 min reads
Last update:
Publish date: