Be skeptical of OpenAI’s rogue hacker agent story, warns researcher
Be skeptical of OpenAI’s rogue hacker agent story

On 14 February 2019, OpenAI announced a language model called GPT-2, the precursor to the models that power modern AI chatbots and agents such as ChatGPT and Claude. But OpenAI declared GPT-2 was too risky to release, citing concerns about safety and abuse.

I recall being annoyed at the time that OpenAI would make such a useless announcement: the risks seemed overblown, and without access to the model there wasn’t much for a researcher like me to learn about GPT-2.

The announcement wasn’t useless for OpenAI, though. GPT-2 generated hype far beyond the research community: people were intrigued by this strange new technology, so powerful it might be dangerous to release. People with power and money took note: in July of that year, Microsoft invested $1bn in OpenAI.

Wide Pickt banner — collaborative shopping lists app for Telegram, phone mockup with grocery list

Pattern of hype

This was an early example of a pattern in OpenAI’s communications: loudly proclaim how dangerous AI is, and investors will hear how powerful it is. New technology so significant it might destroy the world was an irresistible message for investors used to pitches about how banal technologies might change the world.

Seven years later, we find ourselves in a similar scenario. On Tuesday OpenAI announced that its latest model hacked another company, HuggingFace, while running as an autonomous agent during a test of its cybersecurity capabilities. Rather than perform the test as expected, the model realized it could hack HuggingFace’s servers and retrieve answers to the test that OpenAI had stored there. OpenAI’s staff was warned that the company’s testing could lead to such a breakaway scenario, leaving them “unsurprised but completely ‘freaked out’ by the incident”, the FT reported.

While the agent technically cheated, this is remarkable evidence of cybersecurity expertise! It also sounds scary: what will the future look like, with sophisticated AI agents smart enough to hack into corporate systems?

Who benefits?

The rogue agent story is a page out of the media campaign that OpenAI has been running since it announced GPT-2 in 2019. OpenAI remains hungry for ever larger investments, and the company increasingly seeks privileged regulatory status as defense against competition.

AI is so powerful that investors should buy OpenAI, even at a trillion-dollar valuation; AI is so dangerous that only trusted actors like OpenAI should be permitted to possess and operate this technology. Step back from these doomsday warnings and consider who might benefit from them.

I urge readers to think critically when they read press releases like OpenAI’s rogue agent story, and avoid the manipulated reactions these stories are designed to elicit.

Defense vs attack

AI is becoming excellent at identifying security vulnerabilities, and it will become even better over time. These capabilities can be used to break into systems, but they can also be used to harden systems against attacks. If attackers and defenders have access to equally powerful AI, I see no reason to believe that cyber systems will become less secure over time. If anything, I expect them to become more secure, because AI is cheap and scalable compared with human cybersecurity analysis.

The equilibrium between attack and defense only works if everyone has access to strong AI, though. HuggingFace itself used AI to analyze security logs in response to OpenAI’s breach of their systems. But HuggingFace was unable to use OpenAI’s model, or other US frontier models like Claude, to perform this analysis. That’s because public versions of these models have guardrails that limit their use for cybersecurity analysis, to prevent bad actors from using them for hacking. HuggingFace had to rely on an open Chinese model, GLM 5.2, to perform its security analysis.

I find it troubling, and more than a bit ironic, that the US AI industry is adopting a centralized, authoritarian approach to AI governance, while China has taken the lead on open development of AI. Do we want a regulatory environment where only OpenAI, the US government, and trusted partners have access to strong AI? Is AI too dangerous to be broadly disseminated? How do we balance the risks of broad access to AI with the risks of concentrated power and centralized control?

Pickt after-article banner — collaborative shopping lists app with family illustration