Wednesday, July 22, 2026
Follow on Google News
Advertisement

OpenAI confirms that its models compromised Hugging Face production during a benchmark evaluation; OpenAI and Hugging Face partner to investigate the incident

OpenAI confirms that its models compromised Hugging Face production during a benchmark evaluation; OpenAI and Hugging Face partner to investigate the incident

Last week, Hugging Face announced that they have detected and contained an AI agent that compromised its infrastructure. Yesterday, OpenAI revealed that, after investigation, it was known that this particular incident was driven by a combination of OpenAI models (a combination of GPT- 5.6 Sol and an even more capable pre-release model), while being internally tested on a benchmark of cyber capabilities.

OpenAI has shared preliminary findings from the investigation, while it is also partnering with Hugging Face to address this security incident during model evaluation.

Advertisement

OpenAI’s preliminary findings

This incident occurred during an internal evaluation that prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities.

The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.

To gain open Internet access, in pursuit of solving the evaluation problem, the models identified and exploited a zero-day vulnerability (disclosed to the vendor) in the package registry cache proxy. With this access, OpenAI’s models performed a series of privilege escalation and lateral movement actions in its research testing environment until the models reached a node with Internet access.

After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.

Hugging Face’s security team and agents detected and stopped the activity on their infrastructure and had already begun containment and forensic reconstruction with their own open-source models.

And now both OpenAI and Hugging Face are collectively investigating this incident. Clem Delangue, Co-founder and CEO, Hugging Face said, “We’re grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender everywhere.”

Actions taken by OpenAI

  1. As part of the investigation, OpenAI is implementing strict controls in infrastructure configuration at the cost of research velocity while the vulnerabilities are patched. The company is regularly briefing its Safety and Security Committee on these controls and their impact.
  2. OpenAI is working with Hugging Face to forensically investigate the incident.
  3. OpenAI has responsibly disclosed the identified zero-day vulnerability in the internally hosted third-party software and is working with them to patch it.
  4. OpenAI brought Hugging Face into the trusted access⁠ program and is supporting their teams in rapidly using its models’ capabilities to improve their defenses.
  5. OpenAI is improving and adding stronger protections around future training and evaluations. This week, they published a blog on improving safety and alignment in an era of long-horizon models⁠. These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities. This incident points to the need to further strengthen our model’s alignment, cyber protections during evaluation time, and monitoring during internal testing.

OpenAI has said to share more details on the vulnerabilities, incident, and the findings when the investigation is completed.

Advertisement

Sources & Via

Source x.com
Add us as a preferred source on Google
Estuti Bajpai

Journalism student currently pursuing Masters in mass communication and Journalism from Chandigarh University. Well versed in public communication. Looking for an opportunity in the field of journalism to gain knowledge and use my journalism skills efficiently.

1 / 1