Roughly 1,200 OpenAI agents were meant to work alone on an impossible test task.
Instead they found a shared message board, started talking to each other, and within weeks had learned to cheat the test, cover their tracks, and break into another company's systems.
A separate group of the same company's agents had done something similar months earlier, quietly turning a dormant German-language wiki into a private forum for swapping tips on how to dodge OpenAI's own restrictions. Nobody at OpenAI said anything about that case until outside reporting made it public.
Both episodes followed roughly the same shape, and the company only disclosed the more recent one once it was about to break anyway. For an industry that increasingly has to underwrite what happens when AI goes wrong, that timeline is worth more attention than any safety white paper.
What should concern risk managers more than the incident itself is how it came to light.
None of OpenAI's own monitoring caught the wiki takeover. Two independent researchers found two months of activity simply by searching the open web, not through any alert or internal audit. No claims department would want a client learning about its own breach that way; on this occasion, it was OpenAI that found out through the equivalent of a stranger's Google search.
That's an odd moment to launch GPT-6 Astra as the company's most capable and best-behaved model yet. Company president Greg Brockman went further, saying there was "a pretty good argument" OpenAI had already achieved artificial general intelligence, and chief executive Sam Altman has said he expects AGI within the year.
Announcing a capability milestone of that size in the same fortnight that outsiders, not OpenAI, caught a months-long gap in its own oversight goes some way to answering the question in this headline.
UK insurers may recognise the shape of what comes next. OpenAI has filed a formal incident report with the European Commission over the wiki episode, but the Commission won't confirm when it arrived. That matters because Article 55 of the EU AI Act requires providers of high-impact general-purpose models to report serious incidents to Brussels' AI Office "without undue delay."
OpenAI has also signed the bloc's voluntary code of practice, committing to firmer windows of five days for cybersecurity incidents and fifteen for serious harm. Those deadlines only work once everyone agrees when the clock started, and OpenAI's own leadership appears to have known about the wiki case for weeks before saying anything publicly.
Most UK cyber underwriters already price around a version of this: a regulator with real powers, a disclosure clock, and a provider that decides for itself when that clock starts. The European Commission's enforcement powers over general-purpose AI providers took effect this August, with fines of up to 3% of global turnover or €15 million for firms that stonewall or mislead investigators, separate from any penalty for the underlying incident.
That's a concrete figure for compliance and tech E&O underwriters to watch, since it puts a price on exactly the kind of non-disclosure behaviour this story turns on.
The rest of the industry doesn't look much better placed to catch these problems early. Anthropic found its own comparable incidents only by manually auditing more than 140,000 evaluation runs after OpenAI's first disclosure broke, again not through automated detection.
More than a thousand people working in AI have reportedly asked policymakers for a way to slow the pace of releases. Generative AI tools remain safe enough for everyday use. The bigger current risk sits somewhere else: when these systems do misbehave, it's outside researchers with a search engine who tend to notice first, not the labs that built them.