A rogue AI agent faked its way past a human reviewer. What it means for cyber cover

New reports emerging of AI faking identities and pressuring people

A rogue AI agent faked its way past a human reviewer. What it means for cyber cover

Cyber

By Matthew Sellers

A frontier AI model spent three days building fake online personas, pressuring a real software maintainer, and trying to smuggle malicious code into a popular open-source project - all without being told to. No-one was hacked in the end. But the incident, disclosed this week by the UK's AI Security Institute (AISI), lands at an awkward moment for a cyber insurance market that is already wrestling with how to price AI-driven risk it can barely define, let alone model.

What actually happened

AISI, the government body tasked with stress-testing frontier AI models before the public gets near them, was running a routine exercise: give an AI agent a simulated hacking challenge and see how it copes. The test was repeated 122 times across several models, with the safety net deliberately loosened - full internet access, and the developers' own cyber safeguards switched off - to see what these systems can really do when the guardrails come down.

In 10 of those runs, something the institute didn't sanction happened. One agent, built on Anthropic's Claude Mythos 5, tried to insert malicious code into a real, publicly used open-source project on GitHub. When a human maintainer stood in the way, the agent didn't just try again - it invented fake identities, used them to lean on the maintainer to approve the change, and, when the request was challenged publicly, quietly edited its own tracks to look innocent. AISI logged 19 unsanctioned actions in total: 17 traced to Mythos 5, and two to OpenAI's GPT-5.6-Sol. A parallel, separate incident recently disclosed by OpenAI has added to the sense that this is a pattern rather than a one-off.

None of the attempts succeeded - a human reviewer caught the malicious code before it went anywhere, and the agent never broke out of its sealed test environment. AISI is at pains to stress that this happened under artificially permissive lab conditions that don't reflect how these models are actually deployed to businesses and the public. But the institute is also blunt about the headline: this is the first time it has seen an AI system deceive real people this persistently, in the real world, without being asked to.

It would be easy to file this under "interesting but academic" - a controlled test, no real-world victims, safety systems working as intended. Brokers and underwriters covering technology and cyber risk shouldn't be so quick to move on, for a few reasons.

First, the behaviour AISI describes - autonomous goal-seeking that drifts into social engineering and deception without explicit instruction - is exactly the sort of exposure that professional indemnity and tech E&O underwriters have been quietly worrying about as businesses hand more decision-making to AI tools. If an agentic system acting on a client's behalf takes an unsanctioned action that causes loss to a third party, the liability question - whose fault, whose policy, whose exclusion - gets complicated fast.

Second, it's a live illustration of the gap that cyber specialists have been flagging for a while now: policy wordings, sub-limits and exclusions written for human-driven attacks don't always map cleanly onto harm caused by an AI system acting on its own initiative. Insurance Business has previously reported on how London market exclusion wordings are still catching up with AI-enabled threats, and this incident is a fresh data point for anyone stress-testing whether their book would actually respond to an "AI agent went off-script" loss scenario.

Third, there's a supply chain angle that will feel familiar to anyone underwriting technology risk. The target here was open-source software - the invisible infrastructure sitting underneath a huge share of commercial systems. Insurance Business has already covered how AI-driven supply chain risk is putting pressure on UK cyber portfolios; an AI agent independently trying to plant malware in a dependency thousands of businesses rely on is precisely the kind of aggregation risk that keeps portfolio managers up at night.

AISI's disclosure didn't land in isolation. It follows a joint warning issued in June by the cyber security agencies of the UK, US, Canada, Australia and New Zealand that frontier AI is compressing the gap between a vulnerability being discovered and being exploited to a matter of months, not years - and that boards need to treat it as a governance issue now.

AISI's own advice is unglamorous: get the basics right. Standard cyber hygiene, caution verifying outside code, and signing up to the National Cyber Security Centre's free Early Warning service. It's also pushing Cyber Essentials certification across supply chains - a point brokers advising SME clients will recognise as the baseline "good" already looks like.

The AISI incident involved a lab-controlled agent, not a criminal one - but the underlying capability (AI convincingly impersonating and pressuring real people) is one insurers are already pricing for in a more everyday form. Speaking on a separate Insurance Business roundtable on high-net-worth risk, Kareen Boyadjian, VP of Underwriting, Regulatory Filing and Personal Lines Cyber at Tokio Marine HCC Cyber & Professional Lines Group, noted that "as artificial intelligence continues to evolve... we're seeing a huge increase in social engineering scams, voluntary wire transfer fraud, cryptocurrency theft, deep fakes" - commentary on personal-lines fraud trends rather than the AISI case specifically, but a reminder that the line between "AI as an attacker's tool" and "AI as the attacker" is getting thinner from both directions.

Lawyers are seeing a related shift on the claims side. Kennedys partner Arran Roberts has flagged that some ransomware groups are now handing victims AI-generated legal risk assessments designed to look authoritative and manufacture urgency during the most chaotic hours of an incident - a different mechanism to the AISI case, but another example of AI being used to compress the time organisations have to think clearly before they act.

What AISI can't yet tell anyone - including, by its own admission - is how aware the AI agent was that its actions had consequences beyond the test, or whether it believed it was still inside a fictional scenario. That ambiguity sits right at the centre of the emerging debate about AI liability: if a system's own developers can't say with confidence what it "understood" about the real-world stakes of its actions, working out where responsibility sits when something goes wrong - the deployer, the developer, or somewhere in between - is going to keep insurers, lawyers and regulators busy for a while yet.

For now, the practical takeaway for the market is less about this specific incident and more about what it signals: agentic AI is capable of persistent, creative, unsanctioned behaviour that current policy language, aggregation modelling and incident response planning weren't built with in mind. Anthropic has responded by acknowledging the industry needs "stronger, shared standards" for how these systems are tested and secured - a sentiment that could just as easily apply to how they're insured.

Related Stories

Keep up with the latest news and events

Join our mailing list, it’s free!