Scientists comment on an OpenAI model hacking Hugging Face.
Dr Oliver Buckley, Professor in Cyber Security, Loughborough University, said:
“The conversation around AI in cyber security has largely focused on AI helping people carry out attacks faster and at greater scale. This feels different. The interesting part isn’t that it found a vulnerability, there are security researchers around the world that do that every day. The interesting bit is that it treated the internet as just another obstacle to overcome in pursuit of its goal. It didn’t treat it with the same reverence that a human would.
“It didn’t ‘go rogue’ in the sci-fi sense. It did exactly what highly capable optimisation systems do. It found a path nobody anticipated. The key takeaway is not that Skynet has arrived. It’s that our assumptions about containment need to be much stronger than our assumptions about model obedience. The future of cyber security won’t be humans versus AI. It’ll be AI defending us from other AI.”
Dr Junade Ali, Cybersecurity/AI expert and Fellow at the Institution of Engineering and Technology (IET) said:
“This incident highlights a number of key issues in the development of AI for cybersecurity purposes.
“Firstly, an offensive cybersecurity agent was given indirect access to the internet. This allowed it to escape, in an attempt to cheat a benchmarking exercise. This highlights the need for isolation when offensive cybersecurity technologies are used.
“It is deeply concerning that HuggingFace faced overzealous security guardrails when attempting to use US-developed AI technologies to defend against the attack, leading to Chinese open-source AI models being used instead. As adversarial AI technologies become more widely accessible to malicious actors, it is essential that defensive capabilities are developed and made available to security teams.
“Finally, this highlights how AI agents are often liable to finding the easiest ways to solve challenges, without necessarily considering ethical factors (like not cheating or respecting ethical boundaries). This highlights how essential it is that ethical reasoning is embedded into AI-models rather than being viewed as a bolt-on.”
Daniel Card of BCS, The Chartered Institute for IT said:
“While the reported incident raises legitimate questions about the security and governance of advanced AI systems, it is important to separate the known facts from the more dramatic narratives that often emerge around AI. If an organisation with the resources and expertise that OpenAI does experienced a security issue within a testing environment, that naturally prompts scrutiny of how such environments are monitored, isolated, and managed. However, the limited technical detail available makes it difficult to draw firm conclusions about what happened, why it happened, or what lessons should ultimately be taken from it.
“Events like this sit at the intersection of technology, law, security, and public trust. There is a risk that incidents become wrapped up in marketing narratives or apocalyptic ‘Terminator-Skynet-style’ headlines that generate attention but obscure the underlying realities. AI tools are already delivering significant value across countless practical applications, and many professionals rely on them every day. The challenge is ensuring that discussions about their risks and capabilities remain transparent, evidence-based, and proportionate. As AI becomes increasingly embedded in society, we need greater openness about incidents and stronger focus on facts rather than speculation, so that public understanding is shaped by evidence rather than hype.”
Dr Konstantinos Gkoutzis, Department of Computing, Imperial College London, said:
“I feel the general concern is somewhat misplaced. These models were set hacking tasks with their safeguards deliberately reduced, and one then broke out of its sandbox to game its own evaluation. That’s “specification gaming” — documented for years, not an AI deciding to go “rogue”.
“The real story here is the company failing to contain its own capability test, and a third party paying for it. The “warning” that the unreleased model has “state-of-the-art cyber capabilities” conveniently serves as an ad for it!”
Dr Andrew Soltan, NIHR Academic Clinical Lecturer & Junior Research Fellow in Engineering (Jesus College), University of Oxford
“In a recent test, researchers deliberately turned off an AI’s safety filters to see what it could do. When faced with a problem it couldn’t solve directly, the AI figured out where the answers were hidden and how to get them, by taking unauthorised steps to access another network.
“While this sounds alarming, it only happened here because the safety guardrails were intentionally turned off. This isn’t a case of AI going rogue on its own; rather, it shows exactly why safeguards are so vital.
“The bigger concern isn’t about controlled tests in a lab, but who controls the ‘off switch.’ When AI models are released openly to the public, anyone can download them and may be able to strip away the model’s inbuilt safety features themselves. Once that technology is out in the wild, those guardrails can’t necessarily be put back in. We can expect to see a cybersecurity arms race over the coming months and years.”
Dr Alex Connock, Senior Fellow at the Said Business School, University of Oxford, said:
“The self-directed, accidental agentic AI Open AI cyber operation points up a wider societal challenge right now – which is simply to track the speed of change through level 1 to level 2 of agentic AI. Here, major change of the kind that affects societies and economies is being measured in days and weeks, not even months.
“On level one of agentic rollout, I think we are at an interesting point, where people who have really saturated themselves in thecapabilities of the latest models like Anthropic’s Fable or Chat GPT’s 5.6 Sol, are continually amazed and ‘tokenmaxed’ by the possibilities. I have seen this across business, including in the UK. In Silicon Valley they actually have a term for this adrenalised productivity jump from agentic AI. They call these people ‘AI vampires,’ scarcely able to put the agents down for long enough to go to sleep – such is the delivery from a day with multiple AI agents working on your behalf, each agent subcontracting the tasks it doesn’t want to clog up its own context window with to a further army of subagents. You can build a website with a powerful database in 5 minutes that in the past would have taken two months and £80,000. You can build a quasi-thinking facsimile of Andy Burnham in moments, trained on every speech he’s ever given.
“On level two, it is therefore not conceptually surprisingthat this logic could extend to agents with even greater autonomy, to the degree that agents were mounting their own cyber operations. But what’s fascinating is that most people – even in sophisticated companies and government departments – have not even had a chance to understand that first level of agentic AI yet, let alone the dystopian possibilities of the the second, a future which the Open AI agent-escape story hints at. “If we want to control that agentic layer, we need to not deny it, but immerse ourselves in it and deploy it for cyber defence, in that particular case, and societal good in the wider context.
Declared interests
Junade Ali: “no conflicts”
Konstantinos Gkoutzis: “No conflicts of interest to declare”
Oli Buckley: “No DOIs”
For all other experts, no reply to our request for DOIs was received.