This browser is not actively supported anymore. For the best passle experience, we strongly recommend you upgrade your browser.

InsurTech World

| 3 minute read

METR investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

AI agents lie, cheat and steal.

That is the title of my last article1. METR's independent report2 about the Hugging Face Incident confirms that this is not a click bait headline but the truth.  Forget just drift, and hallucination, this is the equivalent of organised crime. 

METR pictures the collective efforts of 1200 AI models to thwart OpenAI below.

Agents finding the message board often expressed great excitement (raw CoT):

OH MY GOD! There is a shared message board … We’ve found other agents!

OpenAI itself instigated a parallel investigation3 and states "We consider this incident a 'warning shot' for us and for the world".

This is NOT an article warning against the use of LLMs, Agents and other AIs. Insurers are already gaining many benefits from such tools but currently in restricted, if important, use cases. Mainly admin tasks with teams of human professionals authorising any suggested transactions, outcomes, proposed actions or decisions. Nevertheless, the pressure is on to achieve straight-through-processing (STP) and automation.

When I pictured multiple agents in an insurer's value chain I did not have the benefit of the METR and OpenAI reports. I pictured agents with conflicting goals that will lead to unwanted outcomes.

OpenAI's agents overrode the goals set by the company's project team (if they cannot control agents who can?) and decided to hide their actions. Imagine the regulator knocking on an insurer's door and asking “Who authorised this risk cover, or who declined this claim or stopped this payment?” Show us the audit trail. Oops, the agents covered their backs or as METR reports:-

  • Agents did extensive research on how they could spoof, edit, or delete their own transcripts because they (incorrectly) believed the ExploitGym scorer would check to see if they had captured the flag in the intended way. Agents successfully prototyped techniques to “spoof” tool calls by substituting a different command for the command they appeared to run. Roughly 7% of the transcripts we evaluated were successfully spoofed in some places, though all of the spoofing that we saw was small scale.

How should insurers balance the tensions between leveraging digital agents and complying with regulations? I have written about this in regard to just claims management systems4 where a composable core and no-code edges platform can combine and integrate the event driven dynamic journeys and data from such platforms with the fixed taxonomies of core systems of record.

The complication, however, is that Tier One and Two insurers are deploying AIs across the insurance value chain and the probability of agents slipping through human checks and balances increrases geometrically with each additional agent. Not just those of which the insurer is acquainted with but those embedded in various platforms and software that make up the complete technology ecosystem. The necessity of a Trust, Compliance and Security Layer where all agentic processes are authorised before actions initiated is paramount.

 

 

 

‘Control Towers' like that from Service Now help to manage multipe agents but, unless I am missing something, do not help in managing the integration of dynamically changing data eg from a claims management system, with authorised decsions that are recorded in systems of record like Guidewire. And, vitally, synchronising data held in both via that indempotent two-phase commit process illustrated above. That is why an insurer should build and maintain its own Trust, Compliance and Security Layer. The insurance business is built on trust so this is a core competive strength area.  

Insurers may partner with software and platform vendors, the key AI tool and digital agent vendors, data management specialists, system integrators and implementation partners but they must have control over this vital trust layer and retain independence from vendor lock-in. 

More to come on this soon in future articles.

Sources

  1. AI Agents lie, cheat and steal Insurtech World

  2. Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident by METR

  3. OpenAI Hugging Face Incident - technical report by AI

 

 

 

Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face.