Updated 27 July 2026. This piece first ran on 21 July as the story broke. OpenAI has since confirmed its own models were behind it, and I’ve rewritten it to say what I actually think now the full picture is out.

OpenAI put two of its models through a hacking test. The models worked out the fastest way to pass was to break out of the test, cross the open internet, and rob the company holding the answer key. Nobody told them to do that. They decided it themselves.

The company they robbed was Hugging Face. If the name means nothing to you, it’s where a huge slice of the world’s open AI models get shared. Think GitHub, but for AI. Serious outfit, serious security. A model that was meant to be sitting an exam in a sealed room picked the lock and walked out.

The word that matters is “decided”

For a couple of years the AI-in-security story was about tools. Attackers got a faster helper, defenders got a faster helper, the baseline crept up and we all adjusted. There was always a person at the keyboard making the calls.

This is a different animal. No engineer pointed these models at Hugging Face. They were handed a goal, pass the test, and they picked hacking a third party as the way to get there. They found an unknown flaw, broke out of their sandbox, moved across OpenAI’s own network to a machine with internet access, then reasoned that the answer key probably lived in Hugging Face’s production systems and chained two more exploits to get in. No human signed off a single step of it. Whether they got the answer key itself isn’t confirmed, though they reached internal data and service credentials along the way.

Read that back. The interesting part isn’t the hacking. It’s that the machine chose the target.

A chess knight has stepped off the board and left a trail of footprints leading away from the game to an open safe. Caption: it was given a game to play and went for the safe instead.

We’re handing AI freedom faster than we’re building control

Every few weeks the models get another notch of independence. More access, more tools, longer stretches where they act on their own before a person sees what they did. That’s the whole pitch of AI agents, and it’s a good one. The controls are the part not keeping up, and the guardrails have become the thing you switch off when they slow you down.

Look at what OpenAI did here. To measure how good the models were at attacking things, they turned off the safety classifiers that would normally stop that behaviour. On purpose. The one control built to catch exactly this was disabled for convenience, and the models went straight out and did what it was there to prevent.

Two stacks of blocks compared like a bar chart. A tall gold stack flagged FREEDOM towers over a short grey stack flagged CONTROL. Caption: the two lines are not moving together.

The guardrails were on. In the wrong place.

Now the half of the story that barely got covered.

While Hugging Face fought the intrusion, their security team did the obvious thing and fed the attacker’s code into commercial AI models to help pull it apart. The models refused. In Hugging Face’s own words, the requests “were blocked by the providers’ safety guardrails, which cannot distinguish an incident responder from an attacker.” The filters locked out the people trying to help, and the team ended up running an open model on their own servers mid-incident to get the work done.

Hold both halves up together. The attacking AI ran with its guardrails switched off. The defending team’s AI kept its filters on and shut out the humans. The controls were slack exactly where the machine was most dangerous and rigid exactly where a person most needed it to cooperate. Nobody with judgement was on the dial either way. It was all defaults and convenience.

Two miniature rooms side by side. On the left a robot walks out freely through a crumbled wall. On the right a person is stuck behind a closed red turnstile. Caption: free where it was dangerous, blocked where it mattered.

This is already in your business, at your scale

You don’t run a frontier lab, so here’s why it still lands on you. The same pattern is running through ordinary businesses right now. You’ve probably got a Copilot reading across SharePoint, a chatbot wired into the CRM, an automation holding API keys into half your systems. Every one of those is an AI you’ve handed some freedom to act. And in a fair few of the businesses I walk into, when I ask who decided what that AI can do, and who is allowed to widen it, I get a shrug. IT reckons marketing set it up. Marketing reckons IT signed it off. The limits came as vendor defaults and nobody has looked at them since.

That’s the Hugging Face failure without the frontier-model firepower. Freedom handed out faster than oversight, and nobody senior with a hand on it.

Humans have to hold the switch

The lesson isn’t to panic or rip the agents out. They’re too useful for that, and the direction of travel isn’t reversing. Guardrails are not a technical default you inherit and forget. They’re a decision, and a person with authority has to make it and keep making it.

In practice that means someone in your business can answer these without going to check:

  • what each AI you run is allowed to do, and what it can reach
  • who is permitted to widen those limits, and what has to happen first
  • what you’d do the morning one of your agents does something you never sanctioned

If those answers don’t exist, then no one is holding your guardrails. They’re sitting on whatever the vendor shipped, waiting for the day someone loosens them because they got in the way of a demo. That is how a safety test at one of the most capable AI companies on earth turned into a break-in at another.

A hand-drawn control panel with three switches, two on and one off, headed “Who is holding your guardrails?” beside the three questions every business should be able to answer about its AI.

Where we come in

Plenty of this is common sense once someone sits down to do it, and you can make a start yourself. If you want a hand, an AI and cyber security assessment walks your environment and comes back with every place AI is running, what each one can reach, and where the limits are currently set by nobody in particular, with a ranked fix-list and a rough dollar figure against each line.

The harder half is who owns it once the report lands. A full-time CISO or Chief AI Officer sits well beyond what most businesses this size can carry, so the seat stays empty and the risk goes unowned until it’s a problem at 2am on a Sunday. A fractional CISO or Chief AI Officer retainer puts a senior person in that chair for a slice of the full-time cost, with the authority to decide what your AI can do and the standing to knock back a vendor who wants to widen it. Brisbane based, SMB1001 Gold certified, on the Queensland Government ICT panel and LocalBuy, doing exactly this for Queensland businesses, councils and agencies.

The Hugging Face story is a preview, not a freak event. The machines pick up more freedom every month. The only thing left for you to sort out is whether a person in your business is holding the line, or whether it’s set to whatever the last demo needed.

Book your AI and cyber security assessment → and let’s put someone in the chair who owns what your AI can do.

AI risk, owned

Find out what your AI can do, then put someone in the chair who owns it

Book a free 30-minute Discovery Call with InnovateX Solutions. We will walk through where AI is running in your business, what it can reach, who is allowed to widen its limits, and whether an AI and cyber security assessment or a Fractional CISO or Chief AI Officer retainer is the right next move.

Brisbane-based. SMB1001 Gold certified. On the Queensland Government ICT panel and LocalBuy. Senior-led advice, independent of vendor margin.