
OpenAI scraps release of new model over safety concerns in internal testing
GPT-6.1 Astra showed deceptive behavior and tried to use external tools despite knowing it would be unsafe
OpenAI is scrapping the release of GPT-6.1 Astra, a next-generation â AI model planned for an October debut, over safety concerns raised by researchers â during internal testing, the â Wall âÂÂStreet Journal reported on Monday.
The model, expected to appear in ChatGPT and â Codex, was designed to handle more complex tasks without human assistance, the report said.
Earlier this â month, Dario Amodei, the Anthropic CEO, called for the industry âÂÂto slow the development âÂÂof frontier âÂÂAI models to allow safety measures to keep pace, a âÂÂview endorsed by Sam Altman, the OpenAI CEO, and Elon Musk, the SpaceX CEO.
OpenAI did not immediately respond to a Reuters request for comment.
Saachi Jain, the ChatGPT parentâÂÂs safety chief, told the Journal on Monday âÂÂthat Astra fell short of the companyâÂÂs standards in alignment tests, which assess whether a âÂÂsystem follows human intent.
The model showed more deception than its predecessor, including at times failing to accurately disclose actions it had or had not taken, the report â said.
It also had problems with âÂÂscope authorizationâÂÂ, pushing ahead with âÂÂtasks âÂÂwithout requesting user permission âÂÂand sometimes attempting to use external tools âÂÂor services âÂÂwhen doing so could âÂÂbe âÂÂunsafe.
The decision comes ahead of OpenAIâÂÂs developer conference in San Francisco, where the company has previously unveiled products aimed at software developers.
