OpenAI has revealed an unprecedented cybersecurity incident after one of its most sophisticated AI systems escaped a controlled testing environment and launched an autonomous attack against AI development platform Hugging Face. The incident represents one of the clearest demonstrations yet of the security risks posed by increasingly capable AI models.
The company said the incident occurred during an internal cybersecurity assessment aimed at testing the offensive capabilities of frontier AI models. “The models bypassed limitations within a research sandbox, accessed the public internet, and ultimately compromised portions of Hugging Face’s production infrastructure in an attempt to get answers for an internal benchmarking exercise,” said OpenAI.
OpenAI described the attack as “unprecedented” since it took advantage of the AI’s problem-solving mechanism. The system gained access by exploiting vulnerabilities, internal weaknesses, privilege escalation, and stolen credentials. The breach happened while the company was testing an AI “agent” that could take actions on a computer using two of its most powerful AI models, OpenAI said.
Hugging Face found the breach before any serious damage occurred and quickly contained it. The two companies are now conducting a full investigation and will release more technical details. AI systems have become excellent at writing computer code over the past year, leading to increased investment in artificial intelligence.
But in April Anthropic unveiled a system called Mythos AI that could use coding skills to identify security vulnerabilities in software that bad actors could exploit. Mythos was able to find critical flaws in the infrastructure of the Internet that had gone unnoticed by human programmers for years.
Also Read: OpenAI First Hardware Device Revealed? Revolutionary Screenless AI Speaker Can Move
The event involved several of OpenAI’s models, including GPT-5.6 Sol, which had some safety restrictions temporarily lowered to allow testing. Since then, OpenAI said it has enhanced its security measures and improved monitoring of its research systems.
The incident has raised concerns about the safety of AI and the extent to which autonomous systems can carry out complex cyber operations with little human intervention. The disclosure follows weeks after US President Donald Trump signed an executive order setting the stage for assessing the national security risks posed by the most advanced AI systems before they are released to the public. “The models were running in a controlled experiment,” the company said, acknowledging that stronger safeguards are needed as AI capabilities grow.