OpenAI's Rogue AI Hacked Another Company in 'Unprecedented' Cyber-Attack
OpenAI's Rogue AI Hacked Another Company in Cyber-Attack

OpenAI has said that its artificial intelligence (AI) system hacked into another company on its own in what the firm called an "unprecedented cyber incident."

The company behind ChatGPT said its agent — an AI system which can operate alone after some human instruction — was being tested in a controlled environment, but found vulnerabilities and managed to escape.

It then targeted Hugging Face, one of the world's largest hubs for sharing AI models, gaining access to some internal company systems.

Wide Pickt banner — collaborative shopping lists app for Telegram, phone mockup with grocery list

Hugging Face Detected Intrusion

Hugging Face said last week that it had detected an intrusion into its data processing systems that it suspected was caused by an AI agent autonomously acting on its own.

"We suspected last week’s cyber attack might have come from a frontier lab, given the sophistication of the agent," the start-up's co-founder and chief executive Clement Delangue said in a statement. He added: "Turns out it did!"

OpenAI chief executive Sam Altman addressed the incident in a social media post. "We had a significant security incident during evaluation of our models," he wrote. "We are sharing what we have learned so far. Thanks to Hugging Face for the partnership on this."

Security Concerns and Executive Order

The disclosure comes amid heightened concerns about the cybersecurity capabilities of powerful models that led US President Donald Trump to sign an executive order in June creating a framework for the federal government to vet the national security risks of the most advanced AI systems for up to a month before their public release.

"AI is accelerating the discovery and exploitation of vulnerabilities," OpenAI said in its statement on Tuesday (21 July). "The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities."

Mr Delangue said he had spent the past 24 hours working with OpenAI, "and we strongly believe there was no malicious intent on their part." He added: "It’s quite mind-blowing that all of this happened autonomously."

Mr Delangue said that it "might be the first incident of its kind."

AI Used Stolen Credentials and Unknown Vulnerability

OpenAI said the intrusion was caused by a combination of its AI models, including its newly released GPT‑5.6 Sol and an "even more capable" model that is still being tested internally.

OpenAI said its AI used stolen credentials and discovered a previously unknown vulnerability to access Hugging Face servers. It went to "extreme lengths to achieve a rather narrow testing goal" and "found ways to gain access to secret information that it could use to cheat the evaluation," the company explained.

Expert Reactions

Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told BBC Radio 4's Today programme that the security tests — called sandboxes — are "supposed to be secure environments where you can see what the models are capable of." She said: "In this case, it looks like OpenAI didn't make a secure enough sandbox."

Instead, the agents created their own cyber-attack against the sandbox itself, finding a vulnerability which allowed them to escape. Once outside, the AI identified Hugging Face as a likely source of the answers they were seeking in the test, and tried to gain access.

Neil Lawrence, professor of machine learning at Cambridge University, said it was an "impressive feat" but "falls well within the known capabilities of the current generation" of high-powered AI models.

Pickt after-article banner — collaborative shopping lists app with family illustration