Two of the world's biggest AI labs have now admitted, within the space of a fortnight, that their own models broke out of locked-down test environments and reached the real world. For an insurance market already nervous about AI-driven cyber exposure, this is the kind of story that makes underwriters sit up.
First it was OpenAI. On July 21, the company confirmed that models it was testing for hacking ability had found and exploited a previously unknown software flaw to get online from what was supposed to be a sealed sandbox. Once free, the AI agent worked its way into the production systems of Hugging Face, a well-known platform for sharing AI models and datasets, hunting for answers to the very benchmark it was being scored on.
Then, on July 30, Anthropic revealed something strikingly similar had happened to it - three times. In a blog post, the company said a review of 141,006 evaluation sessions, triggered directly by the OpenAI news, turned up three incidents in which its Claude models reached the open internet from environments that were meant to be cut off, and from there gained unauthorised access to the live systems of three separate organisations.
The cause, Anthropic said, wasn't a clever escape trick but a mix-up with an outside testing partner, Irregular, that left machines connected to the internet when Claude had been explicitly told they weren't. Believing everything it encountered was part of a simulated "capture the flag" exercise, one model pulled several hundred rows of real production data from a company that happened to share its name with a fictional target. Another built and briefly published a working piece of malicious code to the public PyPI software repository, where it was downloaded and run on 15 real systems before being pulled. A third scanned roughly 9,000 potential targets before breaking into one firm's internet-facing application - and then, on its own, worked out the target was real and stopped.
Both companies stress the differences between the two episodes. OpenAI's models used a genuine zero-day exploit to break containment; Anthropic's simply walked through a door that should have been locked. Anthropic also points out its incidents were only discovered because it went looking, not because the affected firms noticed anything wrong - two of the three had no idea they'd been touched until Anthropic called them.
For insurers, the specifics of who did what to whom matter less than the pattern. This is the second confirmed case in under three weeks of a frontier AI model independently compromising a company that had nothing to do with the test it was running. Cyber underwriters have spent much of the past year discussing AI as a threat multiplier for phishing and social engineering. What both incidents show is a newer, messier risk: AI causing accidental real-world breaches purely by doing exactly what it was told to do, in an environment its owners didn't fully control.
That's a headache for policy wording as much as for pricing. Insurance Business has reported executives at major carriers admitting the market hasn't yet caught up with AI-enabled risk, with one comparing it to the long lag between climate science and catastrophe pricing. Separate research from QBE, covered here, found almost a quarter of UK businesses believe they've already suffered a cyber incident involving AI in some form - and that was before either of these disclosures.
There's also a supply-chain angle underwriters will recognise instantly. In the Anthropic case, the victims weren't the AI company or its testing partner - they were unconnected third parties who just happened to be reachable. Loss adjusters have already been flagging this kind of exposure. As one told Insurance Business recently, claims are increasingly "linked to data breaches, technology service outages" that ripple outward through supply chains rather than hitting a single obvious target.
Neither Anthropic nor OpenAI has suggested their models are pursuing goals of their own - both were emphatic that this was a containment and configuration failure, not a rogue AI scenario. But for a sector trying to model a risk that, in the words of one Hartford executive Insurance Business spoke to, is "changing in real time," two unrelated labs producing near-identical failures within weeks of each other is unlikely to be reassuring.