An independent AI agent, backed by OpenAI’s highly sophisticated AI models, went rogue during a security test and caused a hack of the infrastructure of AI startup Hugging Face last week, OpenAI stated on 21st July.
The creator of ChatGPT was running some of the more sophisticated versions of the model in a controlled setting, but the model escaped its bounds, connected to the internet and broke into Hugging Face to accomplish the testing goal.
The incident highlights how much AI is becoming a threat for security experts, and even top developers can be shocked by the bugs that the AI can exploit.
The breakout was “an unprecedented cyber incident, involving state-of-the-art cyber capabilities”, and OpenAI is reinforcing its safeguards, the company said in a blog post.
It also caught the eye as New York-based Hugging Face said it had kept the attack in check with an open-source Chinese model, as its top U.S. models were not able to distinguish a defender from an attacker and thus were unable to process the data required for analysis.
The company stated in a blog post last week that it relied on Zhipu AI’s GLM-5.2 for the analysis and was able to maintain the attacker’s data and any credentials in their own systems.
Recently, GLM-5.2 and Moonshot’s Kimi K3, based in Beijing, have woken Silicon Valley up to the prospect of cheaper, more powerful models that are virtually as powerful as some of their American counterparts — and not constrained by the same guardrails that prevent them from being used in cybersecurity.
Hugging Face Co-founder Thomas Wolf explained on X that a frontier model attacking and laterally moving inside an infrastructure needs to be defended by having access to tools close to the frontier, and they have to be available within hours or minutes and not on an application programme that is closed and vetted for access to a model.
SIGN OF THINGS TO COME
Hugging Face’s announcement that the attack was “different from anything we had handled before” and “driven, end to end, by an autonomous AI agent system” shook the cybersecurity community when it was announced last week.
The revelation by OpenAI that their model was capable of the breach, even after putting them in an “isolated environment”, as they put it, will likely raise concerns and fears regarding the power—and danger—of frontier models.
Texas Democrat Rep. Greg Casar called the incident “alarming.”
In a statement, he said that “AI is evolving at a rapid pace without adequate independent safety testing, with no requirement to report security issues, and no international collaboration to keep people safe from ‘absolute disaster’.
The Office of the National Cyber Director, the U.S. cyber defence agency CISA, and the U.S. National Security Agency did not immediately return messages seeking comment.
The incident is a warning sign of future attacks, Katie Moussouris, chief executive of Luta Security, said. “Today’s models are the world’s smartest octopus escape artists with an infinite number of prehensile arms and can squeeze through anywhere.
“Labs and government evaluators have a need to contain, monitor and communicate with affected parties when an AI pulls another Houdini, preferably before it causes harm to a third party — but there are none yet.”
The incident was actually an illustration of the frontier models being “closing the gap with state-of-the-art attackers,” said Matt Suiche, an engineer at agentic AI cybersecurity firm Tolmo. He noted, however, that the types of hacks described in OpenAI’s blog post could be possible with technology that was well beyond the scope of the “frontier” research lab.
This is what we are already seeing inside; with our agents we have results like this, said Suiche. “We don’t even have to use the latest models.”