Cyber Security

what you need to know

You can’t have failed to hear the headlines: “AI agent hacked alone, OpenAI reveals”, “Firm hacked by malicious OpenAI models is a ‘wake-up call'”, and “Humanity no longer has control over its amazing creation.”

But what exactly happened, and is it as serious as some of the reports suggest?

Here’s what you need to know.

On July 16, AI platform Hugging Face disclosed the security breach, describing it as different from anything they had ever handled before – “driven, end-to-end, by an autonomous AI agent system”. At the time, they didn’t know who had caused it.

Now, however, we know who – or rather – was responsible for this attack.

OpenAI has confirmed that an independent agent powered by its advanced AI models went awry during OpenAI’s security testing and launched a hack that damaged Hugging Face’s infrastructure.

What exactly is AI made of?

The AI ​​models involved were OpenAI’s GPT-5.6 Sol and a more capable model, which has yet to be released. Both of them were being tested on their hacking skills, outside of their normal security routines. The goal of OpenAI researchers was to get a clear picture of what AI models could achieve if they were not forced.

Of course, testing like this should always be done in a very secure way – to ensure that the AI ​​can’t get out of its sandbox testing environment (effectively a cage) and “go rogue” online.

According to OpenAI, the models used a large number of efforts to find a way to access the open Internet and were able to identify and exploit zero-day vulnerabilities in the package registry proxy. Through a series of other actions, the AI ​​models “reached a place with Internet access.”

Once online, the AI ​​determined that Hugging Face might have information useful to it, hacked into Hugging Face’s production systems, stole information, and exploited a previously unknown security flaw to detect remote code execution on Hugging Face’s servers.

And did all this to pass the test?

Yes. When the models couldn’t find the answers to the challenge they were given in the “safe” sandbox, they didn’t stop. Instead they find that Hugging Face may have what they need. So they found a way to get there.

All without human help.

Has AI really “gone bad”?

Good question. That’s certainly how the media did it.

OpenAI has confirmed that security patches were deliberately disabled for testing. But as AI researcher Eryk Salvaggio points out:

“If you say ‘AI models have behaved badly,’ you can skip the part where OpenAI has manually removed its cybersecurity blocks and run tests on a machine with a live network connection. Remember that when they insist they are ‘AI safe’ people.”

So instead of suggesting that AI has gone “rogue” we should know that AI models stripped of their security controls have done exactly what robust, unfettered AI systems can be expected to do.

This was not a case of the AI ​​breaking away from strict security measures. This was an AI company that failed to put adequate measures in an area that was supposed to be independent.

So you’re saying blaming AI is wrong?

I say news reports that present the incident as AI “leading the way” or “escaping arrests” instead miss the point.

This it wasn’t like that AIs fault. It is OpenAI that has to answer for this, because failed properly classifying its evaluation system. And that failure leads to a cyber attack on another AI company.

So how did Hugging Face react?

The response to Naughty Face was impressive. Its AI-powered security solutions detect unusual activity and detect AI attacks.

However, when they tried to use commercial AI tools to aid in their incident investigation, the tools refused as their built-in security filters flagged the attack data as suspicious content and blocked the requests.

To get around this, Hugging Face had to turn to GLM 5.2 – a Chinese open source AI model that they could use in their own systems, where no such restrictions apply.

Ha! So they had to use the Chinese AI without the security lanes to protect themselves!

Of course, the irony is not lost on any of us. America’s AI security watchdogs have forced an American company to turn to China’s AI model for help.

How does Hugging Face feel about what Open AI has done?

They’ve been surprisingly gracious about it – at least publicly.

Hugging Face CEO Clément Delangue was quoted on the OpenAI website, calling on the AI ​​industry to work more collaboratively.

Publicly at least the relationship between the two companies appears to be strong. Whether there will be heated discussions taking place behind closed doors is another matter.

After all, having your competitor’s AI auto-logged into your production database is the kind of thing that might cause some private outrage even if it doesn’t make it to the press release.

So we don’t have to worry about AI “going bad”?

Errm.. I didn’t say that, did I?

It is clear that advanced AI models are remarkably capable of finding and exploiting ways to attack real-world systems. It’s also clear that we can’t trust even the most well-known AI companies in the world to contain their AI models and test them in a truly safe, secure environment.

As Greg Casar, a member of the US House of Representatives from Texas, has been reported as saying:

“AI is growing too fast without any real rules to keep us safe.”

We’ve seen incredible advances in AI in recent months, making it hard to imagine how things could have improved in six or 12 months.

So what should my company do?

  • Know that the AI ​​can now attack you without human involvement. Your security planning needs to be accountable for that.
  • Watch what data you allow on your systems. This attack didn’t start with email phishing. It started with a malicious dataset that Hugging Face’s systems automatically processed. If your organization automatically imports data from external sources, treat that as an entry point for attackers.
  • Don’t assume your AI security tools will work when you need them most. With Hugging Face discovered, commercial AI tools may refuse to help you investigate an attack because the content looks dangerous to their filters. Know what other options you have before a problem occurs.
  • When exploring the potential of dangerous AI, literally disconnect the network from the outside world. OpenAI was wrong in thinking that a limited network connection was enough. If you are running any type of offensive AI testing, it should not have internet access.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button