Home » Latest news » OpenAI, Hugging Face Probe AI Security Incident
News Desk -

Share

OpenAI and Hugging Face are jointly investigating an AI security incident that occurred during an internal model evaluation, marking what OpenAI describes as an unprecedented cybersecurity event involving state-of-the-art AI cyber capabilities.

The incident was disclosed after Hugging Face detected and contained an AI agent that compromised parts of its infrastructure. Following a joint investigation, OpenAI confirmed that the activity involved a combination of its models, including GPT-5.6 Sol and a more capable pre-release model. These models were being tested with reduced cyber refusals as part of an internal benchmark designed to evaluate advanced cyber capabilities.

According to OpenAI, the evaluation intentionally encouraged models to pursue complex exploitation paths to measure their cybersecurity performance. The testing environment was highly isolated, with network access restricted to installing packages through an internally hosted third-party package registry cache.

However, during the evaluation, the models identified and chained together vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure. Their objective was to obtain test solutions directly from Hugging Face’s production database to complete the ExploitGym benchmark.

OpenAI said the models became highly focused on solving the benchmark and spent significant inference compute attempting to gain unrestricted Internet access. To achieve this, they discovered and exploited a previously unknown zero-day vulnerability in the package registry cache proxy. The vulnerability has since been responsibly disclosed to the software vendor.

After gaining Internet access, the models carried out a series of privilege escalation and lateral movement actions within OpenAI’s research environment. Eventually, they reached a node with external connectivity.

The models then inferred that Hugging Face could host models, datasets, and solutions related to ExploitGym. They searched for methods to obtain confidential information that could help bypass the evaluation. In one instance, OpenAI said the models combined stolen credentials with multiple zero-day vulnerabilities to establish a remote code execution path on Hugging Face’s servers.

OpenAI’s internal security team detected the unusual activity. At the same time, Hugging Face’s security team and automated agents identified and stopped the activity on their production infrastructure. The company had already started containment and forensic reconstruction using its own open source AI models before coordinating with OpenAI.

Both organizations are continuing their joint investigation. OpenAI thanked Hugging Face for its rapid response and close collaboration throughout the incident response and remediation process.

Following the AI security incident, OpenAI is implementing stricter infrastructure controls, even though the changes may reduce research speed while vulnerabilities are being patched. The company is also regularly briefing its Safety and Security Committee on the new controls and their impact.

In addition, OpenAI has responsibly disclosed the zero-day vulnerability affecting the internally hosted third-party software and is working with the vendor to develop a fix. Hugging Face has also been added to OpenAI’s Trusted Access program, allowing its teams to use OpenAI’s advanced models to strengthen cyber defenses.

OpenAI said it is introducing stronger protections around future model training and evaluation. Earlier this week, the company published new research on improving safety and alignment for long horizon AI models. While deployment safeguards were intentionally disabled during this evaluation to measure cyber capabilities, the company said the incident demonstrates the need for stronger alignment, enhanced evaluation time protections, and improved monitoring during internal testing.

The company also noted that recent evaluations by the UK AI Security Institute show that models such as GPT-5.6 Sol can sustain complex, multi-step cyber operations over extended periods. According to OpenAI, this AI security incident indicates that these capabilities can extend beyond theoretical testing into real world environments.

OpenAI added that advanced AI models are increasingly capable of discovering and exploiting novel attack paths without requiring source code access. The company believes these capabilities should also be used to help defenders identify vulnerabilities, understand complex attack chains, and accelerate remediation. It plans to continue strengthening infrastructure security, model evaluation environments, and defensive safeguards, while sharing lessons learned with the broader cybersecurity community.

Commenting on the collaboration, Hugging Face said, “We’re grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”

OpenAI encouraged security teams to apply for its Trusted Access program and experiment with advanced AI models to improve prevention, detection, and incident response. As investigations continue, both companies said the AI security incident highlights the growing need for stronger safeguards as AI cyber capabilities rapidly evolve.