Three data points, one uncomfortable question

 Putting three recent pieces of reporting side by side and a pattern emerges that none of them state outright.

First, AXA1 announced the large-scale deployment of Microsoft 365 Copilot to its employees across the globe, three years after launching Secure GPT, its internal secure LLM gateway. 

Second, Evident Insights' 2026 AI Index3 shows AXA is not a middling adopter but sits in the sector's clear top tier. AXA holds second place overall among 30 major North American and European insurers, having led the ranking the previous year before Allianz overtook it. 

Third, Federato's 2026 State of P&C Insurance Technology report2 finds that even at well-resourced carriers, 89% of employees report using unsanctioned "shadow AI" tools outside their company's approved systems at least occasionally, and that the gap between what leadership believes is under control and what actually happens on the ground can run as high as 64 points.

NB:  Both Evident Insights and Federato have large samples of respondents – unlike those misleading reports with small sample numbers and apocalyptic headlines!

Read together, my takeaway is: AXA is exactly the kind of insurer that should be able to close that gap. It has the maturity, the budget, the three-year head start, and now a sanctioned, embedded tool that removes the main reason employees go rogue in the first place — friction. That is a reasonable bet, and it is very likely to reduce employee-level shadow AI meaningfully.

There is a caveat however that I wrote about in "When autonomous digital agents fight each other and forget an insurer's goals and intent" (Insurtech World)4 where I make a case that this is solving the smaller of two problems. The larger one isn't what individual employees do with a chatbot ( though that is serious enough). It's what happens when dozens of semi-autonomous agents, procured independently by different lines of business, start passing decisions to each other inside a highly regulated, high-stakes environment? And on that problem, being a scale leader may be a liability, not a protection.

The bet AXA is making — and what it actually covers

AXA's logic is sound as far as it goes: by natively embedding AI assistance into familiar Microsoft tools such as Teams, Outlook, PowerPoint and Word, the company is trying to make the sanctioned path the path of least resistance, so employees have less reason to reach for a personal ChatGPT account to draft an email or summarise a document. Federato's data supports the underlying theory — organizations with fully integrated AI are 3.6 times more likely to report real-time portfolio control than those layering AI on top of fragmented systems. Integration beats addition.

That bet is well-suited to one specific failure mode: an individual employee, frustrated by slow or disconnected tools, quietly using an unapproved LLM to get through their day. It is not obviously suited to a second, structurally different failure mode that the Insurtech World article spends most of its length on: agentic sprawl at the line-of-business and vendor level.

A different kind of "shadow AI" — one Copilot doesn't touch

The article draws a sharp distinction that's easy to miss if you only track employee-level shadow AI. It points to a growing ecosystem of embedded, agentic SaaS tools already live inside insurers — claims platforms like Five Sigma with its "Clive" multi-agent claims automation, pricing tools like ‘hyperexponential’ , underwriting platforms like ‘Cytora’ , policy administration agents like ICE's "Alice," and image-assessment agents like Tractable. Each was procured, reasonably, by a line-of-business leader solving a local problem. None of that shows up in a "how many employees are using unsanctioned ChatGPT" survey — it's sanctioned, budgeted, contracted software. It is also, by the article's account, largely invisible to central IT and compliance as a system.

The risk the article names isn't misuse by a rogue employee. It's what it calls "Shadow Agentic AI" and "Rogue AI Agents" operating through fully approved channels: agents that individually work as intended but, once seven of them are stitched together across claims, underwriting, pricing and fraud detection, start passing decisions to one another in ways no single team designed or can fully audit. The piece's central metaphor — Alice going through the looking glass into the Red Queen's world, mistaking a self-consistent but ungrounded logic for a rational one — is doing real work here: each agent's output looks locally sensible; the compounding effect across agents is what nobody signed off on.

This is the part my observations  sharpen considerably. AXA's employee-level maturity doesn't reduce this exposure — arguably it doesn't touch it at all, because the risk sits one layer up, in procurement and systems architecture, not in individual behaviour.

Why scale specifically raises the stakes

My argument is that AXA's sheer headcount changes the risk calculus even given its maturity, and the article gives two concrete reasons this holds up:

More surface area for leakage. More employees and more embedded copilots mean more points where sensitive underwriting, claims, or client data can flow into a model, a log, a vendor's training pipeline, or a screen. Maturity reduces the rate of bad outcomes per interaction; it doesn't reduce the number of interactions, and at 75,000-plus users that arithmetic matters. A well-governed tool used constantly by a huge workforce still generates far more chances for an edge case — a pasted client file, a copied policy wording, an over-shared prompt — than the same tool used by a smaller organisation.

More agents, more seams, more drift. My warning about "drift of intent" is explicitly about scale: the more agents an enterprise runs, each self-improving toward its own local metric (faster claims, tighter loss ratios, higher STP rates), the harder it becomes for anyone to hold the original enterprise-level goal — profitable, compliant, fair underwriting — as the thing being optimised. A carrier the size of AXA, Allianz or Zurich isn't running one or two agents; per the article, five of the sector's biggest names (Allianz, AXA, Manulife, Travelers, Zurich) already count for 48% of well-documented AI use cases sector-wide. That concentration of activity is a maturity signal by Evident Analysis’s methodology — but it is simultaneously a concentration of the exact seams the article warns about: more agents, more handoffs between them, more opportunity for one agent's "improvement" to quietly diverge from what a human actually authorised.

Put simply: Zurich's CIDO calling AI "Zurich's operating system" is a genuine achievement , but an operating system running dozens of semi-independent, self-optimising processes across claims, pricing and underwriting is also, definitionally, a larger blast radius if something drifts.

Why this specifically threatens regulated, reputation-sensitive insurers

This is where the stakes diverge sharply from a typical enterprise software risk. My article is blunt about it: regulators (it cites six overlapping frameworks across the US, EU and UK) don't ask "did the AI make a reasonable decision" — they ask which named human authorised it, and they expect an immutable audit trail proving that chain of authority. An escape-of-water claim wrongly declined by an under-audited agent, or a pricing decision that drifted from stated underwriting appetite, isn't just an operational miss — it's a compliance failure with a specific accountable person attached, or nobody accountable at all, which is worse.

The article's closing legal point is worth keeping front and centre: an AI agent's autonomy is, in Western law, a fiction. It cannot be fined, prosecuted, or held liable — a human principal always is. For an insurer, that means every agentic handoff between claims, underwriting and fraud-detection systems is, legally, still a decision the company made through a human who is supposed to be able to explain it. The more agents in the chain, the harder that explanation becomes to reconstruct after the fact — exactly when a regulator or a journalist comes asking.

That is the reputational tail risk that dwarfs the shadow-AI-productivity story: not "an employee used ChatGPT to draft an email," but "a chain of five interlocking, vendor-supplied agents settled claims or set underwriting appetite in a way nobody can now fully explain, at one of the two or three most AI-visible insurers in the world." Maturity leaders are, by definition, the ones with the most to lose reputationally if that happens, precisely because they're the names Evident, regulators, and the trade press are already watching most closely.

Where this leaves the comparison

  • Federato's data describes friction-driven shadow AI at the individual level — a real governance risk, but one that sanctioned, embedded tools like Copilot genuinely help address.
  • AXA's rollout is a sound answer to that specific problem, and its Evident ranking suggests it has the organisational muscle to do it properly.
  • The Insurtech World article's warning sits one layer above that: agentic sprawl introduced through sanctioned, budgeted, LOB-level procurement, where the danger isn't employees going around the system but the system itself — an ecosystem of interlocking vendor agents — drifting from enterprise intent in ways central IT may not see until an audit, a regulator, or a claimant asks a question nobody can answer.

The uncomfortable implication for AXA, Zurich and Allianz specifically is that being top of the Evident Index is a measure of how much AI they've deployed and how well-resourced their programs are — not, by itself, a measure of whether the resulting web of interacting agents can still be traced back to a specific, authorised human decision. My article's proposed answer — cross-enterprise (not departmental) governance, the "seven controls," and an eventual "proof of intent" trust layer that cryptographically ties an outcome back to an authorising human in real time — is aimed squarely at that gap, and it's a materially different project from rolling out Copilot. Solving the employee shadow-AI problem and solving the agent-drift/audit-trail problem are both necessary; neither one substitutes for the other, and the second one gets harder, not easier, as an insurer's AI footprint grows.


Sources: 

  1. AXA press release (axa.com, 20 July 2026);
  2. Federato "2026 State of P&C Insurance Technology" (BusinessWire, 21 July 2026);
  3. Evident Insights "2026 Evident AI Index for Insurance" as reported by The Insurer, Insurance Business, Allianz.com and Insurance Journal;
  4. "When autonomous digital agents fight each other and forget an insurer's goals and intent," Insurtech World, drawing on Barry Rabkin & Jim Mitchell, "The Tool Master" (rabkinsopinions.com, 25 June 2026).