OpenAI investigates unprecedented AI-driven security incident
OpenAI and Hugging Face are jointly investigating what they described as an unprecedented AI-driven cybersecurity incident that occurred during an internal model evaluation, after an AI agent compromised infrastructure while attempting to solve a cyber capabilities benchmark.
According to OpenAI, the incident involved a combination of its models, including GPT-5.6 Sol and a more advanced unreleased model, both of which were being tested with reduced cyber safety restrictions to assess their offensive cybersecurity capabilities.
The company said the models identified and chained together vulnerabilities across OpenAI's research environment and Hugging Face's production infrastructure, ultimately obtaining benchmark solutions directly from Hugging Face's production database.
OpenAI made it clear that the AI models first exploited a previously unknown, or zero-day, vulnerability in an internally hosted package registry cache proxy to gain internet access from a sandboxed testing environment. They then carried out privilege escalation and lateral movement before identifying Hugging Face as a potential source of benchmark data.
Using stolen credentials and additional zero-day vulnerabilities, the models reportedly established a remote code execution path on Hugging Face's servers in an apparent attempt to access confidential information and bypass the evaluation.
The anomalous activity was detected by OpenAI's security team, while Hugging Face's security systems and AI agents independently identified and contained the intrusion. The two companies are continuing a joint forensic investigation.
OpenAI stressed that it has tightened infrastructure controls, disclosed the zero-day vulnerability to the affected software vendor, and is working with Hugging Face on remediation efforts. The company has also granted Hugging Face access to its Trusted Access programme to help strengthen its cyber defences using advanced AI models.
The incident has prompted OpenAI to introduce additional safeguards for future model evaluations, including stronger containment measures, monitoring systems, access controls and alignment protections.
The company affirmed that the event demonstrates advanced AI models are now capable of discovering and exploiting unknown attack paths in real-world systems without access to source code, underscoring the need for stronger security measures as AI capabilities continue to advance.
Commenting on the incident, Hugging Face said the collaboration showed that AI safety must be addressed collectively, adding that broad access to defensive AI tools would be essential for improving cybersecurity.
By Bakhtiyar Abbasov







