But there's more to the story-
AI systems starting to surpass humans in cyber-security tasks means that “only very, very few cyber experts in the world” could properly understand the way the breach occurred, says Alex Meinke of Apollo Research, an AI-safety group in London.
I had only just published an article about insurers deploying AI at scale and the dangers of agents acting outside the intended rules, security and human oversight perceived to be in place when I saw this BBC article.
" The ChatGPT-maker said its agent - an AI system which can operate alone after some human instruction – was being tested in a controlled environment, but found vulnerabilities and managed to escape.
They targeted Hugging Face, one of the world's largest hubs for sharing AI models, gaining access to some internal company systems……
Once outside, the AI identified Hugging Face as a likely source of the answers they were seeking in the test, and tried to gain access."
Reading OpenAi's own announcement adds context.
Open AI explains that it “This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. ”
Brian Siemsen, the CEO of Wilbur a claims management platform adds:
"OpenAI didn't lose control of a guarded agent. They ran the eval with the cyber refusals turned down on purpose, to see how far the models would push. The agent chained real vulnerabilities and reached its goal. That's not a guarded system failing. That's an unconstrained one behaving exactly as you'd expect.
So the lesson isn't "secure your sandboxes". The capability is already here. The only thing sitting between task-focused and task-obsessed is the constraint layer you decide to keep.
That's the argument for sequence. Evidence first. The human keeps the decision, doesn't hand it over. Automation last, and inside limits you can defend to a regulator.
The agent being clever was never the risk. Giving it the decision before you've built the layer that keeps it honest, that's the risk."
That is the only way to plan, resource and train employees to leverage AI responsibly and safely. Ensure authorised humans have set the rules by which any agents make a decsion and ensure that any agent acting outside those is referred to a human professional with the abilityt and authority to validate or reject the transaction/process.
I wrote about what can happen if you don't in 'When automous digital agents fight each other and forget an insurer's goals and intent'.
Picture this where different agents have conflicting goals just as it appears that the ChatGPT agent prioritised tyeb goal of finding an answer, scaled the digital prison walls of the sandbox and then sought to worm its way into Hugging Face's secure data .
One key leeson to apply is that when experimenting in sandboxes make sure they are secure and do not hand over from humans to agents to automate workflows and decisions until you have measurable evidence the constraints on agents are absolute and enforced.
OpenAI states “We’ve responsibly disclosed the identified zero-day vulnerability in the internally-hosted third-party software and are working with them to patch.” It's not just your own systems, it's all tghe third party software in youur technology stacks. Teke thisd announcementb to a software vendor for example:

What happens if youi missed the announcement; when yoou first licensed t5he software or core system it did not have these AI features. How do the changes impact your own platforms?
OpenAI did not expect the flaw and did not know about it until Hugging Face detected the attacks . In one example, the agent model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers.
Do as OpenAI have and “strengthen your model’s alignment, cyber protections during evaluation time, and monitoring during internal testing.”
The whole incident strengthens the validity of the article I wrote -
Scale Is Not the Same as Safety: What AXA's Copilot Rollout Reveals About the Real Risk in Insurance AI
OpenAI said the incident was "unprecedented", external, and it was conducting an investigation alongside Hugging Face, whose boss Clement Delangue said in a post on X it was "mind-blowing that all of this happened autonomously".
unknownx500