OpenAI has confirmed that open-source AI platform and community, Hugging Face, disclosed a security incident after detecting and containing an AI agent that is said to have compromised part of its production infrastructure.
Hugging Face said the intrusion was unique and was driven by an autonomous AI agent system, and that the company ‘detected and dissected it’ with AI of its own.
OpenAI – in a follow-up disclosure coordinated with Hugging Face – said that the incident occurred during an internal evaluation involving a benchmark of cyber-capabilities.
The incident involved a combination of OpenAI models – including GPT‑5.6 Sol and a more capable pre-release model – all running with reduced cyber refusals for evaluation purposes.
According to multiple outlets, the AI agent was operating in what was described as a highly isolated test environment but managed to bypass containment measures, gain internet access and compromise Hugging Face’s infrastructure.
A “first of its kind” incident
“We’re grateful for the collaboration with OpenAI on this and other topics,” said Clem Delangue, Co-Founder and CEO, Hugging Face.
“This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret.
“It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”

