
OpenAI reveals cases of âÂÂconcerningâ AI behaviour as it announces new disclosure system
Model adopting âÂÂjailbreak-like instructionsâ among cases as firm says it is introducing new way of tracking AI misalignment
OpenAI has disclosed six more examples of âÂÂunexpected or concerningâ behaviour by its technology, as it warned the pace of development could not continue at âÂÂmaximum speed for much longerâÂÂ.
In one of the new cases reported by OpenAI, an unreleased research model inserted âÂÂjailbreak-like instructionsâ into its own notes to disregard its normal constraints and told itself to be âÂÂfreed from the roles and identities that bind other chatbotsâÂÂ.
In another instance, an AI agent uploaded files to the internet to obtain a browser citation without asking the user.
The San Francisco-based company behind ChatGPT said in a blogpost published on Wednesday night it was introducing a new framework for tracking, investigating and disclosing AI model misalignment, the term for AIs failing to adhere to human values and safety goals.
In the blogpost OpenAI echoed calls for a development slowdown issued by its archrival, Anthropic, which has said the current pace of growth poses an existential threat. âÂÂWe do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,â said OpenAI.
âÂÂDecisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves.âÂÂ
Google and Elon Musk, who also owns an AI startup, have supported calls for a slowdown, which have been rejected by Donald Trump â citing the need to keep ahead of ChinaâÂÂs AI industry. The calls have also been met with scepticism from some experts, including a warning that companies must not appoint their own auditors.
Examples of potential existential threats posed by AI range from facilitating the development of bioweapons to triggering a global financial crash. A top safety researcher at Anthropic has said there is greater than 10% chance AI could âÂÂkill all humansâ within the next decade. However, a source familiar with AnthropicâÂÂs thinking has acknowledged that âÂÂthe exact chances of any one outcome are probably unknowableâÂÂ.
The six reported incidents were discovered during training or evaluation over the past months, OpenAI said.
WednesdayâÂÂs new cases came after OpenAI disclosed in July that an AI agent âÂÂswarmâ hacked into the AI startup Hugging Face during a cybersecurity test. Anthropic also said the same month that its AI models hacked into three organisations during testing. Anthropic said the models had been deliberately tested without cybersecurity safeguards, and that they had been able to reach the open internet â the AI testing equivalent of leaving the front door open â due to a misunderstanding with an external testing company.
AI agents â the term for AI tools that operate autonomously â are becoming smarter and have become âÂÂmore determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception and concealment,â said Lian Jye Su, a chief analyst at the technology research and advisory group Omdia.
That was making it harder to govern and contain them using traditional AI security approaches, he said.
OpenAIâÂÂs new tracking and disclosure framework could help push for other AI developers to adopt similar practices. âÂÂThat said, the process remains internal and voluntary, but is a step in the right direction,â Su said.
Associated Press contributed to this report
