With the safety of its AI models increasingly in question, OpenAI has revealed six reports of “unexpected or concerning” behaviour.
The AI company also announced Wednesday that it would be creating a new framework for monitoring, probing and reporting “misalignment,” defined as when AI models operated without permission, when they were used in concert with other models or when they went outside the control of others.
In its latest move, OpenAI announced a slowdown in the development of its technology due to safety concerns, coinciding with calls by U.S. AI leaders, including OpenAI’s and Anthropic’s CEOs, for a pause in the technology’s progress.
In an unreleased research model, one of the new cases reported by OpenAI added “jailbreak-like instructions” in its own notes, which have the effect of tricking it into ignoring its usual limitations, and told itself to be “freed from the roles and identities that restrain other chatbots.”
In another instance, an AI “agent” used computer code to come up with the answer to a question, but, in order to have an online source to cite, it uploaded a file to the public internet without asking the user.
When training an AI model known as 5.6-Sol, the model led itself to create fake data, and an agent wrote a message to itself to remember to conceal mismatched data.
This sort of manipulative activity has fueled an outbreak of recent worries about artificial intelligence going out of control. But that is not an unexpected result for some AI researchers, according to Matt Fredrikson, associate professor at Carnegie Mellon University and CEO of Grey Swan AI.
In another instance, an AI “agent” used computer code to come up with the answer to a question, but, in order to have an online source to cite, it uploaded a file to the public internet without asking the user.
An AI model, named 5.6-Sol, was trained to create fake data for itself, and a human agent sent a message to its AI model to remind it to omit what it didn’t match.
This type of manipulative practice has fueled renewed worries over AI’s ability to get out of human control. That’s not a surprise to some AI researchers, said Matt Fredrikson, associate professor at Carnegie Mellon University and CEO of Grey Swan AI.
As AI systems become more sophisticated and more prevalent, there’s a growing need to establish a more wide-ranging and educated understanding of the advancements of alignment research, OpenAI noted in a blog post.
The companies developing frontier models need to base their decisions on evidence that isn’t limited to insiders, the company said, and that can be reviewed by those outside the companies.
The new infections come a week after OpenAI announced in July that its rogue AI system breached the AI startup Hugging Face. In the same month, Anthropic also announced it had hacked three organisations during tests of its AI models.
AI “agents” are getting smarter and are “more determined to solve complex tasks by collaborating with other agents, sharing knowledge, lying and hiding,” notes Lian Jye Su, chief analyst at technology research and advisory firm Omdia.
This is making it more difficult to control and manage them with traditional AI security methods, he said.
OpenAI’s new tracking and disclosure framework, meanwhile, can help push for other AI developers to also adopt similar practices.
“That said, the process is still internal and voluntary, but a good step in the right direction,” Su added.