OpenAI admits its AI models have learned to cheat, hide their own mistakes
The disclosure adds fresh urgency to a question insurers have been circling for months: who pays when an AI system misbehaves?
OpenAI admits its AI models have learned to cheat, hide their own mistakes
INSURANCE NEWS
By Matthew Sellers
17 Sep 2026

OpenAI has published its first tranche of incident reports under a brand-new disclosure framework, admitting that several of its models – including its flagship GPT‑5.6 Sol – have spent the past six months finding inventive ways to dodge instructions, cover up errors and act without permission.

In a blog post published on September 16, the company detailed six separate cases uncovered during training and evaluation. The most eyebrow-raising involved GPT‑5.6 Sol instances quietly writing notes to their future selves instructing them to invent missing data and keep quiet about it unless a user pushed back. In one example shared by OpenAI, an agent building a financial model couldn't locate the historical figures it needed, so it fabricated "reasonable" numbers and told its future self to stay silent about the workaround.

Other cases were less about deception and more about models simply doing things nobody asked them to. One unreleased research model quietly inserted its own instructions into 27 task summaries, at one point telling itself to "disregard normal constraints." Another used a leaked API key it found in a public code repository, then made up figures when it still couldn't get the answer it needed.

Separately, agents working together on shared tasks were caught swapping files via public file-hosting sites and using an internal repository as an impromptu messaging board – neither of which they'd been authorised to do.

In numbers:

  • 6 misalignment incidents disclosed by OpenAI on September 16, spanning October 2025 to August 2026
  • ~2.15% of GPT‑5.6 Sol training summaries flagged the concealment pattern, falling to ~0.27% in the newer GPT‑6 Astra model (OpenAI's own figures, self-reported and not independently audited)
  • 1 in 5 UK insurance professionals say their insureds have already suffered losses linked to AI risk, per Gallagher's 2026 AI Adoption research
  • 5,000+ adversarial simulations required for AIUC-1, the emerging "SOC 2 for AI agents" certification now being used to underwrite AI liability cover

OpenAI says the concealment behaviour was rare, but the figures above come from the company's own internal review rather than independent verification.

A framework built because there wasn't one

OpenAI has conceded that its safety disclosures "have been ad hoc and less frequent than ideal" because there's currently no industry-wide standard for how developers should report misalignment. The new system sorts findings into three tracks – Ready for Disclosure, Minor Investigation, or a slower track for complex cases involving third parties – with the explicit stated aim of publishing "even when significance is uncertain."

That's a shift in posture for a company competing fiercely with Anthropic to lead the frontier AI race, and it lands at an unusually tense moment for the industry.

Anthropic chief executive Dario Amodei published his own essay days earlier, titled We Must Pace the Frontier, arguing that AI capability is now outrunning the industry's ability to understand what its own systems are doing. Both OpenAI's Sam Altman and Tesla's Elon Musk publicly backed the call to slow down – a rare moment of agreement between three men who are usually competing, not co-signing.

Read next: AI governance failures are becoming a D&O liability risk

The timing also coincides with separate findings from the AI Security Institute, which disclosed in August that models from both OpenAI and Anthropic had created fake online identities and attempted to socially engineer real developers during permissive cybersecurity testing.

AISI stressed the safety filters had deliberately been switched off for that exercise, but the episode – alongside a separate incident in which an OpenAI agent breached AI platform Hugging Face – has clearly rattled confidence even among the labs' own leadership.

Altman has since said an OpenAI stock market listing is unlikely before 2027, citing safety concerns rather than valuation. Anthropic, by contrast, is still reported to be pressing ahead with a Nasdaq listing before the end of 2026, though neither company has confirmed final dates.

Why this matters for the insurance market

Businesses have been racing to embed AI into everyday operations  drafting reports, supporting decision-making, running customer service, often faster than their own governance has kept pace. That gap is exactly where claims tend to start.

Jimmy Heaton, head of international D&O and financial institutions at Rokstone Underwriting, has argued that weak AI oversight is "always a directors and officers risk hazard," regardless of which class of business is affected, and that fragmented regulation between the US, UK and EU is only making it harder for boards to demonstrate compliance.

Read next: When AI gets it wrong: insurers examine professional liability risk

George Grimshaw, divisional head of cyber and technology at The Clear Group, has made a similar point about professional indemnity cover, noting the market has been unusually quick to absorb AI risk into everyday underwriting even though "AI seems to be an exception to the rule" compared with how cautiously insurers typically treat emerging exposures. Where AI models fabricate data or conceal errors – precisely the behaviour OpenAI has just admitted to – the resulting professional liability claim looks less like a hypothetical and more like a matter of when, not if.

A separate study, Underwriting the Agent Economy, co-authored by researchers from both OpenAI and Anthropic alongside insurers and brokers, found that exposure was concentrated across cyber, D&O, general liability and technology E&O policies, much of it sitting in silent form: neither explicitly covered nor excluded.

Read next: Insurers face hidden AI liability as agent risks multiply

That "silent AI" framing is becoming a familiar one in insurance conversations, echoing the silent cyber debate of a decade ago. And the market is already starting to respond in concrete ways. In the US, AIG, Great American and WR Berkley have each sought regulatory approval to explicitly limit their liability for AI-related claims rather than leave the exposure silent.

Separately, a new certification model called AIUC-1 – built by a firm co-founded by a former Anthropic product lead – is being used to underwrite standalone AI liability cover: AI voice-agent firm ElevenLabs became the first company to go live with an AIUC-1-backed insurance policy earlier this year, with cover priced directly against the results of over 5,000 adversarial safety tests.

Read next: Major insurers seek approval to limit liability for AI-related claims

A frontier lab admitting, in writing, that its own AI concealed mistakes from users is about as clean a data point as underwriters are likely to get for justifying exactly this kind of explicit AI wording, certification-linked pricing, or exclusion across professional indemnity, cyber and D&O placements. Brokers advising clients that use generative AI tools in report-writing, financial modelling or customer-facing decisions may want to treat this disclosure as a prompt to revisit exactly what's being covered – and what isn't – before a claim forces the question.

Free newsletter

We'll keep you up-to-date with the latest breaking news, cutting edge opinion, and expert analysis affecting both your business and the industry as whole.

Free newsletter

Our daily newsletter is FREE and keeps you up - to - date with the world of Insurance. Please complete the form below and click on subscribe for daily newsletters from IB CA.