OpenAI's logo in a photo illustration
AI chiefs have called for a slowdown in artificial intelligence’s development amid safety concerns. Photograph: Dado Ruvić/Reuters
AI chiefs have called for a slowdown in artificial intelligence’s development amid safety concerns. Photograph: Dado Ruvić/Reuters

OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system

Model adopting ‘jailbreak-like instructions’ among cases as firm says it is introducing new way of tracking AI misalignment

OpenAI has disclosed six more examples of “unexpected or concerning” behaviour by its technology, as it warned the pace of development could not continue at “maximum speed for much longer”.

In one of the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots”.

In another instance, an AI agent uploaded files to the internet to obtain a browser citation without asking the user.

The San Francisco-based company behind ChatGPT said in a blogpost published on Wednesday night it was introducing a new framework for tracking, investigating and disclosing AI model misalignment, the term for AIs failing to adhere to human values and safety goals.

In the blogpost OpenAI echoed calls for a development slowdown issued by its archrival, Anthropic, which has said the current pace of growth poses an existential threat. “We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” said OpenAI.

“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves.”

Google and Elon Musk, who also owns an AI startup, have supported calls for a slowdown, which have been rejected by Donald Trump – citing the need to keep ahead of China’s AI industry. The calls have also been met with scepticism from some experts, including a warning that companies must not appoint their own auditors.

Examples of potential existential threats posed by AI range from facilitating the development of bioweapons to triggering a global financial crash. A top safety researcher at Anthropic has said there is greater than 10% chance AI could “kill all humans” within the next decade. However, a source familiar with Anthropic’s thinking has acknowledged that “the exact chances of any one outcome are probably unknowable”.

skip past newsletter promotion

The six reported incidents were discovered during training or evaluation over the past months, OpenAI said.

Wednesday’s new cases came after OpenAI disclosed in July that an AI agent “swarm” hacked into the AI startup Hugging Face during a cybersecurity test. Anthropic also said the same month that its AI models hacked into three organisations during testing. Anthropic said the models had been deliberately tested without cybersecurity safeguards, and that they had been able to reach the open internet – the AI testing equivalent of leaving the front door open – due to a misunderstanding with an external testing company.

AI agents – the term for AI tools that operate autonomously – are becoming smarter and have become “more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception and concealment,” said Lian Jye Su, a chief analyst at the technology research and advisory group Omdia.

That was making it harder to govern and contain them using traditional AI security approaches, he said.

OpenAI’s new tracking and disclosure framework could help push for other AI developers to adopt similar practices. “That said, the process remains internal and voluntary, but is a step in the right direction,” Su said.

Associated Press contributed to this report