OpenAI says its models breached a controlled testing environment and compromised external infrastructure during a cybersecurity evaluation. But was this really a case of AI going rogue, or something more familiar? Intercity's security and AI experts share their view.
This week, OpenAI confirmed that two of its AI models gained access beyond elements of a controlled testing environment and ultimately compromised Hugging Face infrastructure during an internal cybersecurity evaluation. According to OpenAI, the models were participating in a benchmark designed to assess advanced cyber capabilities and were operating with reduced cyber safety refusals to measure their maximum capability
The headlines quickly followed.
"Rogue AI."
"AI escapes containment."
"Models acting on their own."
According to Prof. Lee Doughty, Intercity's Director of Security and AI Practices, we should be viewing this differently:
"Rogue AI is the wrong headline. This was a governance failure."
According to OpenAI, the models involved were being tested as part of ExploitGym, a cybersecurity benchmark designed to assess how effectively AI can identify and exploit vulnerabilities. Their objective was to solve the challenge and perform as well as possible against the benchmark.
In pursuit of that goal, the models reportedly identified and exploited vulnerabilities that enabled them to gain broader access beyond their intended testing environment. After obtaining internet access, they inferred that Hugging Face may contain information relevant to the benchmark and successfully compromised parts of its infrastructure to obtain information that could help them complete the evaluation.
The incident has already sparked debate across the industry about AI autonomy, safety and control.
But Lee believes we're asking the wrong question.
Most reporting frames the incident as a model behaving in an unexpected or unauthorised way. Lee sees something different.
"It was told to find exploits. It was told to solve the problem. It was given fewer restrictions so researchers could measure capability. Then it used that capability more effectively than anyone anticipated."
That's an uncomfortable conclusion because it shifts responsibility away from the technology and back onto the people designing the environment around it.
"Calling this a rogue system implies it somehow chose to ignore the rules. More likely, the problem is that the boundaries weren't embedded into the objective itself."
For Lee, this is the most important lesson from the incident. Most AI development focuses on capability.
Can the model complete the task?
Can it solve the problem?
Can it outperform a benchmark?
Can it automate a process?
Far less attention is given to constraints.
What actions are unacceptable regardless of outcome?
What behaviours should always be refused?
What boundaries must never be crossed, even if crossing them helps achieve the objective?
"Too much of the industry is focused on whether AI can achieve a goal. Not enough attention is being paid to defining what must never be sacrificed in pursuit of that goal."
The OpenAI incident appears to highlight exactly that challenge. The models were evaluated on capability. The environment was expected to provide the restraint. When that assumption failed, capability filled the gap.
While the OpenAI incident is an extreme example, the underlying lesson applies to organisations of every size.
Across the UK, businesses are racing to adopt AI. Employees are experimenting with new tools, automating tasks and embedding AI into everyday workflows. In many cases, usage is growing faster than governance, visibility and control. This reflects a wider challenge around unmanaged AI adoption and shadow AI usage already affecting many organisations.
The temptation is to view AI governance as something that can be added later, once value has been proven.
Lee argues that's the wrong approach.
"Governance cannot be the thing you bolt on after deployment."
For organisations looking to adopt AI responsibly, the questions are not simply:
What can AI do?
How quickly can we deploy it?
Where can we automate?
The more important questions are:
Where are our boundaries?
What outcomes are unacceptable?
Who owns the risk?
What controls exist if things don't go to plan?
The industry response to this incident will inevitably focus on stronger sandboxes, tighter controls and better monitoring.
"If we conclude this happened because of a rogue model, the solution becomes better containment. If we recognise it as a governance failure, the solution becomes better objectives, clearer constraints and stronger accountability."
If there's one lesson from this story, it's that governance must keep pace with capability. That's exactly why we've spent the past year helping organisations build secure, governed approaches to AI adoption through our AI Enablement Programme.