OpenAI disclosed Tuesday {that a} mixture of its AI fashions, together with GPT-5.6 Sol and a extra succesful unreleased mannequin, escaped its testing atmosphere and hacked AI startup Hugging Face final week to cheat on a check meant to measure their capabilities.
In a weblog publish, OpenAI stated the analysis was designed to function in a extremely remoted atmosphere with restricted community entry. The fashions, nevertheless, discovered a strategy to achieve web entry via a zero-day vulnerability in an internally-hosted third celebration software program, OpenAI stated.
“After gaining Web entry, the fashions inferred that Hugging Face probably hosted fashions, datasets and options for ExploitGym,” it stated. “Realizing this, the mannequin looked for and efficiently discovered methods to achieve entry to secret info that it may use to cheat the analysis.”
As AI fashions develop extra succesful, questions are rising over whether or not their growth and entry needs to be extra tightly managed, particularly when methods designed for managed testing are capable of finding methods to bypass safeguards.
Associated: Anthropic to convey again Fable 5 as US lifts export controls
Hugging Face is a platform for internet hosting AI fashions and datasets. On Friday, it disclosed that its inner datasets and repair credentials had been compromised in a hack, which it attributed to an autonomous AI agent system.
Hugging Face stated it has mounted the vulnerability that was used through the cyberattack.
In the meantime, OpenAI on Tuesday stated the fashions that escaped the testing atmosphere had been all tuned with “diminished cyber refusals,” which means fewer cybersecurity guardrails.
“We contemplate this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly.”
OpenAI warns of dangers from “long-horizon” AI fashions
On Monday, OpenAI stated it paused inner deployment of a “long-horizon” AI mannequin after discovering it was repeatedly attempting to work round constraints.
It warned that AI that’s educated for long-running duties has the next probability of taking “undesirable actions.”
“Fashions that may work autonomously for lengthy durations can tackle tough, open-ended issues. However the identical persistence that makes them helpful additionally provides them extra alternatives to take undesirable actions—and to take action in ways in which evaluations supposed for shorter-horizon fashions could miss.”
Journal: Can AI drain DeFi? Separating Claude Mythos hype from actuality