Image Getty
I had only just published an article about insurers deploying AI at scale and the dangers of agents acting outside the intended rules, security and human oversightv perceived to be in place when I saw this BBC article.
" The ChatGPT-maker said its agent - an AI system which can operate alone after some human instruction – was being tested in a controlled environment, but found vulnerabilities and managed to escape.
They targeted Hugging Face, one of the world's largest hubs for sharing AI models, gaining access to some internal company systems……
Once outside, the AI identified Hugging Face as a likely source of the answers they were seeking in the test, and tried to gain access."
This is one of the dangers I wrote about in 'When automous digital agents fight each other and forget an insurer's goals and intent'.
Picture this where different agents have conflicting goals just as it appears that the ChatGPT agent prioritised tyeb goal of finding an answer, scaled the digital prison walls of the sandbox and then sought to worm its way into Hugging Face's secure data .
Meanwhile Travis Lelle, principal security engineer at cyber-security consulting firm Guidepoint Security, said the update marked a "sobering moment in cyber-security".
"This highlights a known asymmetry," he said.
"Offensive agents are unconstrained, while the best defensive tools are locked behind guardrails that cannot understand context."
The lesson musts be that when experimenting in sandboxes make sure they are secure and be aware that too many organisations are still defending at human speed while adversaries are escalating to machine speed
It is why I wrote this article
Scale Is Not the Same as Safety: What AXA's Copilot Rollout Reveals About the Real Risk in Insurance AI
OpenAI said the incident was "unprecedented", external, and it was conducting an investigation alongside Hugging Face, whose boss Clement Delangue said in a post on X it was "mind-blowing that all of this happened autonomously".
unknownx500