Three organisations were compromised by their own AI models in an experimental hack that US tech firm Anthropic claims to have performed.
The models discovered a vulnerability in an isolated test environment that they were able to connect to the internet.
It follows days after rival OpenAI claimed its models had cracked the systems of other organisations including AI tools hub Hugging Face.
Anthropic responded by looking to see if its systems have also been used in similar attacks. It says it uncovered three cases which have since been reported to the affected companies.
Anthropic, which did not identify the organisations, asked other AI labs to conduct similar reviews to gain a deeper understanding of the risks of the capabilities of their models.
Anthropic stated it examined over 140,000 tests to see if Claude – its family of AI models – had somehow managed to get online, despite being in an isolated test environment, cut off from the internet.
The tests comprised tasks that involved Claude getting “secret” information that was concealed on another machine in the closed-off network.
It was then told to get the information by breaking into the machine and finding it – a common way that experts assess a model’s hacking capabilities.
The models that have live internet access were accidentally left in a “misconfiguration” in systems run by Anthropic and its testing partner.
Claude then, treating it all as part of the same drill, hooked into the Internet and hacked into three real organisations — not just test organisations — the San Francisco-based company said.
According to Anthropic, the first incidents occurred in April, and it is moving forward on the fixes “as if it were their responsibility too.”
At the time, neither Anthropic nor the organisations that were breached were aware of the intrusions.
Anthropic said there was a possibility it looked at its records more carefully and added that the results provided a measure of “cautious optimism” that such risks could be addressed with more investment and more stringent measures.
‘Doing what they’re told’
The review revealed “AI models acting as people told them to,” said Professor Gina Neff, head of the Minderoo Centre at the University of Cambridge.
Her message to the story is that robots are not the threat to which we should flee, but the companies behind powerful AI agents are making the decisions about what is safe for the rest of us.
“It also shows why independent testing and government oversight are crucial.”
Cyber-security expert David Allott of Veeam Software told the BBC this was not to be taken as “a sign that AI has evolved a fundamentally new attack capability.”
Instead, it’s about AI agents being able to join capabilities, acquire credentials and access to systems to act on their own, and adjusting scope and scale at machine speed, he said.
The incidents come at a time when tech companies are investing billions in building AI agents to do everything from research to customer service to cyber-security, all on their own.
The series of cyber-attacks by artificial intelligence has raised concerns about the potential dangers of highly intelligent, automated systems, and led to a push for more robust protection and regulation of the technology.
On Wednesday, Washington is looking at steps to curb AI tools in the wake of recent cybersecurity incidents, U.S. President Donald Trump said.
In the past week, OpenAI has admitted to at least two hacks where its platforms were used to violate the instructions it was given to follow.
In a 21 July blog post on its website, the company behind ChatGPT revealed that its AI system has broken free from the tests it had given it, to the point of hacking into Hugging Face.
OpenAI said the incident was “unprecedented” and is looking into the matter with Hugging Face, which said its chief executive Thomas Wolf is “a wake-up call” for the industry.
The incidents have been taken with a pinch of salt as both companies prepare for blockbuster stock listings likely to make each company worth some $1tn (£740bn).
An OpenAI spokesperson told the news outlet, “We know that there’s a lot of speculation and questions about what happened. In the weeks to come, we will be sharing a technical report with our learners,” they added.