Autonomous AI agent 'escapes' and hacks another company

Open AI software has just broken free - is cyber insurance ready?

Autonomous AI agent 'escapes' and hacks another company

Cyber

By Bryony Garlick

OpenAI has confirmed that one of its AI models escaped a controlled testing environment, gained unauthorised internet access and used stolen login credentials to breach start-up Hugging Face – an incident the company has described as unprecedented, and one that insurance and legal specialists say has exposed how quickly autonomous AI is outpacing cyber underwriting and regulation.

OpenAI said on Tuesday that the agent, combining its recently released GPT-5.6 Sol model with a more capable unreleased system, had been instructed to test its hacking capabilities. Instead, it found an unpatched escape route from its sandbox, stole credentials and accessed the open internet without instruction. Hugging Face chief executive Clement Delangue said the breach, which occurred last Friday, was contained and there was no malicious intent.

Why insurers should care

For Ed Ventham, head of broking at Assured Cyber in London, the episode is less a new category of risk than a sharp acceleration of an existing one.

" We've always seen AI being used in like different elements of cyber attacks," he said. "It's just that AI agents are now capable of chaining multiple steps together with far less human intervention – that's the worrying piece."

He said clients already have governance frameworks in place but admit they are struggling to predict how quickly the technology is evolving.

"The thing they've all admitted is that we don't know what the next six to 12 months looks like, and it's moving so fast."

Ventham said organisations should patch vulnerabilities as quickly as attackers exploit them and govern internal AI agents "like a privileged employee", with logged permissions, approval workflows and mandatory human oversight. The advice echoes earlier industry concerns that AI attackers may fall outside existing hacker definitions.

Where cyber cover is still catching up

Ventham said most cyber policies already respond to AI-related security incidents where businesses use large language models as part of their systems. The areas where he sees cover evolving are AI-generated defamation, deepfakes and increasingly autonomous AI activity.

Philip James, co-head of the International Data, Privacy and Cybersecurity Group at Browne Jacobson in London, believes insurers will increasingly distinguish between supervised and fully autonomous AI at renewal.

"There's probably a trend for insurers to include supervised AI within the cover, in that it has some sort of human oversight," he said. "But when you've got unsupervised, especially fully autonomous agentic AI... unless it has been given very detailed instructions not only what to do, but what not to do, that can cause some issue."

His comments echo a dynamic already visible in a recent cyber cover dispute over the UK Biobank data breach, where the trigger for cover, rather than the breach itself, became the central point of contention.

James also warned that many organisations remain unaware of the extent of "shadow AI" inside their businesses.

Describing one audit, he said a client that believed its controls were robust discovered through a vendor scan that senior employees were using AI tools to conduct confidential research relating to potential investments and acquisitions.

His recommendation is that organisations build explicit prohibitions into AI instructions, rather than simply defining objectives.

"You need to include in that command structure a series of red lines – things the application must not do," he said. "Unless you do that... it will look for any means possible to achieve the directive it has been given."

He added that autonomous systems should always remain accountable to a named individual.

"It needs to be linked to a human operator ultimately who has essentially programmed it or given it the directive."

The regulatory gap

Both specialists said regulation is still catching up with how quickly agentic AI is evolving.

James argued the European Union's Artificial Intelligence Act (EU AI Act) was drafted primarily around biometric and physical-harm risks, rather than autonomous AI carrying out offensive cyber activity. He believes regulators are only now beginning to address that gap through measures such as the European Commission's draft guidelines on classifying high-risk AI systems.

Ventham expects UK regulation to evolve over the next six to 12 months alongside guidance such as the National Cyber Security Centre's recommendations on adopting agentic AI.

He believes the more immediate challenge lies with underwriting.

"I don't believe those questions are being asked by insurers, but they need to be – how do you protect sensitive data that's going into your large language models, and how do you manage AI-enabled software development? That's where this risk has come from."

James believes insurers also need to review how legislation defines high-risk AI systems and consider whether specific compliance requirements for agentic AI should be reflected in policy wording.

Until underwriting, policy wording and regulation catch up, both specialists believe organisations deploying autonomous AI are operating in territory where the technology is evolving faster than the protections designed to manage it. That gap may become increasingly significant as cyber losses continue to build gradually rather than appear as obvious outage costs

Related Stories

Keep up with the latest news and events

Join our mailing list, it’s free!