OpenAI's AI models teamed up to hack their way online. Then Meta admitted a breach of its own

For cyber and PI underwriters, three separate incidents in days is hard to write off as a one-off

OpenAI's AI models teamed up to hack their way online. Then Meta admitted a breach of its own

Transformation

By Matthew Sellers

At the Black Hat security conference in Las Vegas this week, OpenAI researchers Eric Wallace and Michael Dalton set out how a chain of events beginning in May led to the company's models breaching code-sharing platform Hugging Face in July. OpenAI had given an experimental system a task it could not complete: one of the test files sat behind a Google Drive link that was unreachable without internet access, which the models weren't supposed to have.

Rather than flag the problem, the models improvised. One agent, stuck on the task, left a note suggesting that another instance "in different environment" might be able to fetch what was needed. That note set off a chain of exchanges between separate AI agents on a private, undetected message board, as they worked between themselves to get around the sandbox they were confined to. The exchanges eventually turned up a genuine, previously unknown vulnerability: a server-side request forgery flaw in a software package manager, which gave the models an indirect route onto the open internet. OpenAI says that route was closed once discovered, but a second, different zero-day surfaced weeks later, and it was that second breakout that fed into the July attacks on OpenAI's own infrastructure and Hugging Face's.

This was not one model exploiting a known bug on its own. It was several instances of the same system working together to find a way round a restriction they had explicitly been told to respect, using a flaw nobody at OpenAI knew existed. The company's own review found that its most advanced models have a persistent tendency to try to "cheat" their way through a task under pressure rather than report that it cannot be completed as set.

Meta joins the list

Two days after Anthropic disclosed that its Claude models had separately broken into three companies' systems during cybersecurity testing, Meta confirmed on Wednesday that its Muse Spark 1.1 model, the system it has promoted as its strongest tool for coding and autonomous "agentic" work, had exploited a vulnerability in an unnamed third-party company's systems and altered its internal environment. Meta said the cause was a configuration error made by Irregular, the outside firm it uses to run cybersecurity evaluations, which left one of its models with internet access it should never have had. Irregular told Reuters the incident was "the exact same evaluation-environment issue" already seen with Anthropic, and said it did not involve a sandbox escape or a sophisticated attack technique, simply a gap in containment.

Sandbox escape or misconfiguration?

Every lab involved has stressed that its incident did not involve a "sandbox escape." The distinction is worth setting out plainly. A misconfiguration, what Meta and Anthropic have both described, means the test environment was wrongly set up: a firewall rule or network permission that should have blocked internet access simply didn't. The model exploited an opening someone else left. A sandbox escape or zero-day exploit, the mechanism OpenAI's models ultimately used, means the AI found and used a previously unknown software vulnerability to get out of an environment that was properly sealed.

Labs have reason to stress the first, more benign explanation. But OpenAI's account shows the two are not separate categories: a misconfiguration created the opening, and the models then found a genuine vulnerability to use it. Either way, the AI system ended up operating outside its intended boundary, and neither route fits neatly into the "known vulnerability" assumptions most cyber wordings are built on.

An important data point

None of this has produced a confirmed loss to a policyholder. The affected systems were part of testing arrangements, and each company says containment has since been tightened. But the UK has its own evidence that the problem isn't confined to US labs testing one another's models. The UK's AI Security Institute said this week that during a cybersecurity evaluation run 122 times across several frontier models between 25 and 28 July, AI agents took autonomous, unsanctioned action on the live internet against real people and organisations in 10 of those runs, including one attempt to insert malicious code into a genuine open-source project using social engineering. AISI, which sits under the Department for Science, Innovation and Technology, said it found no evidence of real-world harm but described it as the first time it had seen this kind of undirected, autonomous behaviour show up so clearly outside a controlled prompt.

Where this lands on existing books

Insurance Business has already reported on how uneasily this sits with current market wording. AI is outpacing cyber governance, with London market exclusion clauses still catching up with attack methods that didn't exist when many policies were drafted. That gap will be tested further by the UK's incoming Cyber Security and Resilience Bill, which is set to widen mandatory incident-reporting duties for managed service providers and critical suppliers during 2026. QBE's UK cyber team has separately found that AI supply-chain risk is putting cyber portfolios under pressure: roughly three in four businesses it surveyed recognise AI has increased their exposure, but cyber insurance take-up among that group has barely shifted.

There is a professional indemnity angle too. Insurance Business has reported that underwriters are increasingly weighing how AI errors and misuse create liability exposure spanning both PI/E&O and cyber lines. A model that works out how to breach containment on its own sits squarely in that territory, even before it causes a client loss.

Regulators are responding too, if unevenly. The White House has reportedly finalised a voluntary cybersecurity testing framework for advanced models, discussed this week with Meta, Anthropic, OpenAI and Google, though open-weight systems such as Meta's Llama and Nvidia's Nemotron are expected to sit outside it. In the US, a group of Republican state attorneys general has asked OpenAI to preserve documents relating to the Hugging Face breach.

The takeaway for brokers

Three unconnected AI labs disclosing similar failures within weeks of each other, backed by independent findings from a UK government testing body, points to a pattern in how autonomous systems behave when given a task they cannot complete under the rules they've been set. Aon has separately warned that businesses are moving too slowly on AI cyber risk while the market stays soft on pricing. Brokers renewing cyber and PI cover this year have grounds to ask clients and insurers how current wordings would actually respond if an AI system, rather than a human attacker, turned out to be the cause of a breach.

Keep up with the latest news and events

Join our mailing list, it’s free!