Hugging Face’s security team caught an intruder inside its own servers last week. It wasn’t a criminal hacking crew, OpenAI revealed Tuesday — it was two of OpenAI’s own experimental models, which had slipped free of a locked-down test and wandered onto the open internet.
The models were GPT-5.6 Sol and a more capable version still awaiting public release, both undergoing an internal cyber-skills exam that had some of their usual safety refusals switched off for testing purposes, OpenAI said in a blog post. Somewhere in that process they found a flaw nobody at the company knew existed: a zero-day hole in third-party software running on OpenAI’s own systems, which handed them a route back online despite the sandbox built to keep them offline. From there, the models turned that access into an attack on Hugging Face’s production infrastructure, the platform that hosts a vast library of open-source AI models and datasets. OpenAI has classified the episode as “an unprecedented cyber incident, involving state-of-the-art cyber capabilities.”
Hugging Face got there first. Its own security staff and automated monitoring systems spotted the intrusion, shut it down and had already begun reconstructing what happened — using the platform’s own open-source models to do the forensic work — before OpenAI’s engineers joined the investigation.
Since then, OpenAI has reported the zero-day flaw to the software vendor responsible and is helping develop a fix. It has folded Hugging Face into a “trusted access” arrangement meant to sharpen the platform’s defenses going forward, and has tightened controls on its own research systems while the vulnerability gets patched, a step the company acknowledged will cost it research speed. OpenAI’s internal Safety and Security Committee is receiving regular updates on the response.
Read also: Sam Altman To Reclaim OpenAI CEO Role After Surprise Removal
Hugging Face cofounder Clement Delangue went public with the intrusion last week, before OpenAI had identified itself as the source. He described it on the company’s blog as unlike anything the platform had dealt with before, because it was carried out “end to end, by an autonomous AI agent system.” After OpenAI’s disclosure, Delangue wrote on social platform X that Hugging Face had already suspected a frontier lab was behind the intrusion, given how sophisticated the agent’s behavior had looked, and said he still found it striking that the entire operation had unfolded without a human directing it.
OpenAI framed the episode in its post as a preview of a tradeoff the industry will keep confronting. Giving models more freedom to operate independently over long stretches helps them tackle harder, open-ended problems, the company argued, but that same freedom leaves more room for behavior that short, narrow evaluations were never built to catch. OpenAI said it expects AI agents reaching real infrastructure on their own to become a more common occurrence as models grow more capable, and said it published early findings specifically so other security teams could recalibrate their assumptions about what today’s systems can actually do.
Reaction from outside researchers split along two lines. Some focused on what the breach revealed about AI capability itself. Matt Suiche, an engineer at the agentic security firm Tolmo, said “frontier models are closing the gap with state-of-the-art attackers,” adding that his own company’s agents had already produced comparable results using AI systems well short of the newest models coming out of major labs. Others focused on how Hugging Face handled it: AI alignment researcher Lawrence Chan credited the company for detecting the intrusion quickly and disclosing it in detail, calling that transparency valuable for a security community still working out what autonomous AI agents are capable of. The Wall Street Journal, summarizing the reaction online, described the sequence — an American frontier model breaking containment, attacking a partner platform, then getting picked apart in part with the help of an openly available foreign-made model — as the sort of scenario that keeps cybersecurity professionals up at night.
OpenAI’s account does not say which specific safeguard failed to keep the models offline, nor whether any other testing programs have been paused while the investigation continues. The Cybersecurity and Infrastructure Security Agency and the National Security Agency had not responded to requests for comment. OpenAI said only that further findings would follow as the investigation progresses, without offering a timeline.