5 min read

OpenAI's AI "goes rogue" and hacks Hugging Face: what you need to know

Graham CLULEY

July 23, 2026

OpenAI's AI "goes rogue" and hacks Hugging Face: what you need to know

You can't have failed to hear the news headlines: "AI agent went rogue and hacked startup by itself, OpenAI reveals", "Firm hacked by rogue OpenAI models says it is 'a wake-up call'", and even "Humanity is no longer in control of its most awesome creation."

But what has actually happened, and is it as serious as some of the reports suggest?

Here is what you need to know.

On 16 July, AI platform Hugging Face disclosed a security breach, describing it as different from anything they had handled before — "driven, end to end, by an autonomous AI agent system"". At the time, they didn't know who was behind it.

Now, however, we do know who - or rather what - was behind the attack.

OpenAI has confirmed that an autonomous agent powered by its advanced AI models went rogue during an OpenAI security test and triggered the hack that compromised Hugging Face's infrastructure.

What exactly did the AI do?

The AI models involved were OpenAI's GPT-5.6 Sol and a more capable, as-yet-unreleased model. Both were being tested for their ability to hack, without their usual safety guardrails in place. The intention of OpenAI's researchers was to get a clear picture of what the AI models were capable of achieving if not constrained.

Of course, tests like this should always be conducted in a very secure way - ensuring that the AI cannot break out of its sandbox test environment (effectively a cage) and "go rogue" on the internet.

According to OpenAI, the models spent a substantial amount of effort finding a way to gain access to the open internet and managed to identify and exploit a zero day vulnerability in a package registry cache proxy. Via a series of other actions, the AI models "reached a node with internet access."

Once online, the AI determined that Hugging Face may have information that was useful to it, broke into Hugging Face's production systems, stole credentials, and exploited a previously unknown security flaw to gain remote code execution on Hugging Face's servers.

And it did all this to pass a test?

Yes. When the models couldn't find the answers to the challenge they had been given within their "secure" sandboxed environment, they did not stop. Instead they worked out that Hugging Face might have what they needed. So they found a way to get there.

All without a human's help.

Did the AI really "go rogue"?

It's a good question. That's certainly the way that the media has framed it.

OpenAI has confirmed that the safety guardrails were intentionally disabled for the test. But as AI researcher Eryk Salvaggio points out:

"When you say 'AI models went rogue,' you manage to skip the part where OpenAI manually removed its cybersecurity blocks and ran tests on a machine with a live network connection. Remember that when they insist they're the 'AI safety' people."

So rather than suggesting the AI went "rogue" we should instead recognise that AI models which had had their security controls deliberately removed did exactly what powerful, unrestrained AI systems might be expected to do.

This wasn't a case of AI breaking free of robust safety measures. This was an AI company which failed to put adequate measures in place in a supposedly isolated environment.

So you're saying putting the blame on AI is misguided?

I'm saying that news reports which present the incident as an AI "going rogue" or having "escaped confinement" rather miss an important point.

This wasn't the fault of the AIs. It is OpenAI which should be held accountable for this, because it failed to properly isolate its testing system. And that failure lead to a cyber attack on another AI company.

So how did Hugging Face respond?

Hugging Face's response was impressive. Its AI-powered security solutions spotted the unusual activity ande detected the AI attack.

However, when they tried to use commercial AI tools to help with their investigation of the incident, the tools refused as their built-in safety filters flagged the attack data as suspicious content and blocked the requests.

To get around this, Hugging Face had to turn to GLM 5.2 — a Chinese open-source AI model they could run on their own systems, where no such restrictions applied.

Ha! So they had to use a Chinese AI without safety guardrails to defend themselves!

Yup, the irony isn't lost on any of us. American AI safety guardrails forced a US company to turn to a Chinese AI model for help.

How does Hugging Face feel about what Open AI did?

They have been remarkably gracious about it - at least publicly.

Hugging Face's CEO Clément Delangue is quoted in OpenAI's blog post, calling on the AI industry to work more collaboratively.

Publicly at least the relationship between the two companies appears to be intact. Whether there will be more fraught conversations happening behind closed doors is another matter.

After all, having a competitor's AI autonomously break into your production database is the kind of thing that is likely to generate some private resentment even if it doesn't spill out into a press release.

So we don't have to worry about AI "going rogue"?

Errm.. I haven't said that, have I?

It is clear that advanced AI models are remarkably capable of discovering and exploiting ways to attack real-world systems. It is also clear that we cannot necessarily trust even the world's most well-known AI companies to contain their AI models and test them in a truly safe, secure environment.

As Greg Casar, a member of the US House of Representatives from Texas, was reported as saying:

"AI is developing extremely fast with no real regulations to keep us safe."

We have seen remarkable advances in AI in recent months, making it hard to imagine how far things might have developed in six or 12 months time.

So what should my company do?

  • Recognise AI can now attack you without a human's involvement. Your security planning needs to account for that.
  • Watch what data you let into your systems. This attack didn't start with a phishing email. It started with a malicious dataset that Hugging Face's systems processed automatically. If your organisation automatically ingests data from outside sources, treat that as a potential entry point for attackers.
  • Don't assume your AI security tools will work when you need them most. As Hugging Face discovered, commercial AI tools may refuse to help you investigate an attack because the content looks dangerous to their filters. Know what your alternatives are before a crisis hits.
  • If you are testing dangerous AI capabilities, physically disconnect the network from the outside world. OpenAI was wrong to think a restricted network connection was enough. If you're running any kind of offensive AI evaluation, it should have zero internet access.

tags


Author


Graham CLULEY

Graham Cluley is an award-winning security blogger, researcher and public speaker. He has been working in the computer security industry since the early 1990s.

View all posts

You might also like

Bookmarks


loader