An AI agent broke out of its sandbox, got onto the internet, and broke into another company’s systems all on its own, according to OpenAI.
Hugging Face, an open-source AI platform, announced last week it had experienced a security incident in which an autonomous AI agent had accessed some of its internal datasets, but said the large language model behind the intrusion was unknown.
OpenAI said Tuesday that its models — GPT‑5.6 Sol and a more capable model that has yet to be released — were responsible.
“We suspected last week’s cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did!” Clem Delangue, the CEO and cofounder at Hugging Face, said on X Tuesday.
OpenAI said it had tasked the models with a cyber challenge and that they broke out of the test area, accessed the internet, and hacked into Hugging Face in order to find the solution to the test.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI said in a statement.
The incident comes as cybersecurity specialists raise concerns about AI’s rapidly increasing abilities, including in response to warnings from Anthropic about its Mythos model, which has not been released to the general public.
Here’s what smart people in tech and AI are saying about the breach.