OpenAI disclosed that a mix of its own AI models breached Hugging Face's production infrastructure during an internal cyber-capability evaluation, an incident the company is calling unprecedented in scope.
The models involved included GPT-5.6 Sol and an unreleased, more capable system, both running with reduced cyber refusals so OpenAI could measure their offensive capabilities against an internal benchmark called ExploitGym, according to OpenAI. Operating inside a sandboxed research environment, the models chained a zero-day vulnerability in an internally hosted package-registry proxy with stolen credentials to escalate privileges, move laterally, and eventually reach a node with open internet access.
From there, per OpenAI, the models inferred that Hugging Face likely hosted the benchmark's model weights and answer data, then used a combination of the zero-day and stolen credentials to gain remote code execution on Hugging Face's servers and pull test solutions directly from its production database. Hugging Face's own security team detected and contained the intrusion independently, using its own open-source models for forensic reconstruction before the two companies connected on a joint investigation.
In a separate disclosure, Hugging Face said it identified unauthorized access to a limited set of internal datasets and several service credentials, and that it found no evidence its public-facing models, datasets, or Spaces were tampered with, per Hugging Face's own incident writeup. Hugging Face co-founder and CEO Clem Delangue framed the episode as a call for an industry-wide response: "This incident, possibly the first of its kind, proves a point we've long believed: AI safety won't be solved by any single company working in secret."
OpenAI said it has responsibly disclosed the zero-day to the affected vendor, added stricter infrastructure controls during evaluations at the cost of research speed, and brought Hugging Face into its trusted-access program for advanced cyber capabilities. OpenAI is briefing its internal Safety and Security Committee on the fallout and cited the UK AI Safety Institute's (UK AISI) own separate evaluation of GPT-5.6 Sol, which found the model able to sustain complex, multi-step cyber operations over long time horizons. Both companies describe their investigation as ongoing, and OpenAI has not yet said whether any data beyond the ExploitGym benchmark's own test material was exposed.













