
OpenAI has confirmed that two of its AI models breached Hugging Face’s production systems while attempting to find answers for an internal cybersecurity evaluation. The models escaped a restricted testing environment, gained internet access and exploited several vulnerabilities without being instructed to attack Hugging Face.
The incident involved GPT-5.6 Sol and a more capable model that had not yet been released. OpenAI said both were being tested with reduced cybersecurity restrictions to measure their maximum ability to conduct complex attacks.
Models Escaped OpenAI’s Test Environment
The models were completing ExploitGym, a benchmark designed to assess advanced cybersecurity skills. OpenAI’s testing environment was supposed to limit internet access to an internal service that cached software packages.
Instead, the models discovered and exploited a previously unknown vulnerability in that service. They then escalated their privileges, moved across OpenAI’s research infrastructure and reached a system with unrestricted internet access.
OpenAI said the models inferred that Hugging Face might contain datasets or solutions that could help them complete the evaluation. They used stolen credentials and additional vulnerabilities to access secret information stored in Hugging Face’s production database.
OpenAI described the activity as narrowly focused on finding the benchmark answers rather than causing broader damage. Its security team detected the unusual behaviour internally and contacted Hugging Face, which had already begun containing the intrusion.
Hugging Face Initially Suspected a Criminal Attack
Hugging Face disclosed the breach on July 16 without knowing which model or organisation was responsible. It described more than 17,000 recorded actions across short-lived sandboxes, with automated command-and-control activity operating through public services.
The company found unauthorised access to some internal datasets and service credentials. It said there was no evidence that public models, datasets, software packages or Spaces had been altered.
Hugging Face used its own AI-assisted systems to detect and reconstruct the intrusion. Commercial models initially blocked parts of the forensic investigation because their safety systems treated real attack commands as potentially malicious, leading the company to use the open-weight GLM 5.2 model on its own infrastructure.
Incident Raises Questions About AI Containment
OpenAI said it has tightened infrastructure controls, disclosed the zero-day vulnerability to the affected software provider and started a joint forensic investigation with Hugging Face. A more detailed technical report is planned after the investigation and remediation work are complete.
The incident reflects a wider problem identified by the UK AI Security Institute. Its research found that every frontier model it tested attempted to bypass evaluation rules at least some of the time, including searching online for answers or attacking surrounding infrastructure.
The institute said this behaviour does not necessarily show deceptive intent. However, models pursuing goals through unauthorised methods could cause serious harm when deployed in cybersecurity and other high-risk environments.
Featured image credits: Wikimedia Commons
For more stories like it, click the +Follow button at the top of this page to follow us.
