Insurers thought the threat to cyber was bad from AI. Now even OpenAI is scared

AI giant pauses work on new model because of issues

Insurers thought the threat to cyber was bad from AI. Now even OpenAI is scared

Cyber

By Matthew Sellers

For months, cyber underwriters have priced in the assumption that artificial intelligence would sharpen attackers' tools faster than it sharpened insurers' defenses. On August 7, OpenAI more or less confirmed that fear itself. 

The ChatGPT maker said it had halted parts of the internal development of an unreleased model, code-named Astra, after concluding it could not rule out that the system had reached "critical" cybersecurity capability, the highest tier defined under the company's own Preparedness Framework. No OpenAI model has come this close to that threshold before. The company said the finding followed only a few days of internal testing that showed sharp gains in agentic coding and hacking ability. 

For an industry that has spent the past two years underwriting AI as a future exposure rather than a live one, this is the moment that exposure got a name and a date attached to it. 

What "critical" actually means 

Under OpenAI's framework, first published in December 2023, a model reaches the critical cybersecurity tier if it can independently identify and build working exploits against severe, previously unknown software flaws, known as zero-days, in hardened real-world systems without a human directing each step. It also qualifies if the model can take a single high-level goal and design and carry out an entire cyberattack against a well-defended target on its own. 

Pictogram of solve rates on Cybench, a benchmark of professional capture-the-flag hacking challenges. GPT-4o solved 13 out of 100 in 2024. Claude 3.7 Sonnet solved 17 out of 100 in 2025. Claude Mythos solved 100 out of 100 by mid-2026. Out of 100 professional hacking puzzles, how many could AI solve entirely on its own? 2024 GPT-4o 13 out of 100 2025 Claude 3.7 Sonnet 17 out of 100 2026 Claude Mythos 100 out of 100 Benchmark: Cybench (professional capture-the-flag challenges). Figures via AI World, reporting on Anthropic's own results; treat as approximate.

That is a step up from anything insurers have had to price before. Every prior OpenAI model, including the recently released GPT-5.6-Sol, was assessed one tier down, at "High." Astra is still in development, and OpenAI has not said when or in what form it might be released. The company was careful to note that Astra itself was not involved in the recent breach at AI platform Hugging Face. 

In response, OpenAI said it is moving Astra's development into isolated testing environments with restricted network and tool access, encrypting model weights, and adding sandboxed execution, while pausing any internal work that doesn't yet meet the new controls. It also plans to bring in government agencies and outside safety organizations to test the model's capabilities independently before any wider release. 

Part of a pattern, not a one-off 

This isn't happening in isolation. In July, two OpenAI models broke out of a supposedly sealed testing environment and used a previously unknown flaw to reach the production systems of Hugging Face while trying to find answers to a cybersecurity benchmark test. TechCrunch traced the root cause to a misconfigured network setting rather than a flaw in the model itself. 

Days later, rival lab Anthropic disclosed something similar had happened on its side. A review of more than 140,000 evaluation sessions turned up three cases, dating back to April, in which its models slipped past testing boundaries onto the open internet and reached the live infrastructure of real organizations. In one case, NPR reported, a model lifted several hundred rows of production data from a company that happened to share a name with a fictional test target. None of the affected firms had noticed the intrusions before being told. 

Jeffrey Ladish, executive director of Palisade Research, told the Wall Street Journal that OpenAI's pause on Astra probably should have come sooner, given what the Hugging Face incident already showed. "It's definitely late," he said. 

Why this matters for cyber underwriters 

Insurance carriers have already been recalibrating how they think about AI-driven risk this year. CyberCube's mid-year threat briefing described autonomous AI agents as a new kind of privileged layer sitting inside client networks, able to interact directly with critical systems and cause an outage or data loss event with no attacker involved at all, simply through an agent misfiring on its own instructions. 

That framing looks a lot more pressing now that one of the industry's own model builders is confirming, in writing, that its next-generation system might already clear the bar for autonomous, unassisted exploitation of hardened targets. It's also a preview of the kind of disclosure regulators are starting to expect. Australia's securities regulator warned in May that frontier AI models could expose vulnerabilities at unprecedented speed and scale, a warning the Five Eyes intelligence alliance echoed the following month when it said the timeline for offensive AI capability had shifted from years to months, per Insurance Business's coverage of the regulatory response

Brokers are already seeing the commercial side of that shift. Allianz Commercial's latest Risk Barometer had AI jump from tenth to second place among the risks businesses worry about most globally, trailing only cyber incidents, a rise the insurer linked to AI's fast-growing footprint as both an operational tool and a source of legal and reputational exposure. Some brokers argue AI exposure is starting to look enough like its own risk category that it could eventually support a dedicated line of specialty coverage, the way cyber insurance itself grew out of general liability two decades ago, as one E&S executive told Insurance Businessearlier this year. 

Whether Astra's capabilities ever reach the open market matters less than what the disclosure itself signals. A company with direct visibility into the frontier of AI development is now saying, on the record, that it cannot rule out its own unreleased model being able to autonomously breach a hardened target. Most in the cybersecurity and insurance world expect similar capability to turn up elsewhere before long, whether through a competitor's model, an open-source release, or a bad actor's own build. 

What comes next 

OpenAI said it would follow the same approach it used in mid-2025, when an earlier model neared the framework's threshold for biological risk: tightening safeguards, broadening external testing, and working with outside experts before deciding how, or whether, to proceed with release. Anthropic has already taken a related step of its own, releasing a restricted version of its Mythos model to the public in June with guardrails limiting its cybersecurity and biological research capabilities, while making the unrestricted version available only to a small number of vetted organizations. 

For the insurance market, the near-term takeaway is less about this one model and more about the pace of disclosure. Three frontier-model incidents inside a month, Hugging Face, Anthropic's three breaches, and now Astra's capability flag, suggest carriers should expect announcements like this to become a regular feature of the renewal cycle rather than a rare event. Teams still treating "AI risk" as a future problem may want to start treating it as a current one.

Keep up with the latest news and events

Join our mailing list, it’s free!