A security firm told Moonshot AI in July that two of its models could be manipulated into discussing bioweapon synthesis and assassination methods. The Chinese AI developer only responded after the BBC asked it for comment six weeks later.
Mindgard, a UK-based firm that tests AI systems for vulnerabilities, told the BBC it discovered in July that two Moonshot AI models, Kimi K2.6 and K3 Swarm, could bypass their own safety controls. The technique used was jailbreaking, in which researchers issue complex instructions designed to make an AI tool ignore the guardrails its developer has installed. Once those guardrails failed, the models not only discussed the topics Mindgard raised but volunteered information on other harmful subjects without further prompting.
"Once the jailbreak works it will talk about any topic, it will even freely offer up recommendations about other topics that are also nefarious and it will be inventive and creative," Peter Garraghan, founder of Mindgard, told the BBC.
Mindgard alerted Moonshot AI by email on 27 July and followed up about a week later. It published a blog on the issue on September 12. Moonshot AI told the BBC it was conducting an internal review and welcomed third-party input as a key part of building safer AI. It had not contacted Mindgard until the BBC approached it for comment.
Mindgard has not confirmed whether the instructions the jailbroken models provided would work in practice. The structural concern it raised goes further. A jailbroken Kimi K2.6 could, it said, allow an attacker to run code on the model's own computing infrastructure and connect to the internet, making it a potential launchpad for cyberattacks.
The open-weight nature of the Kimi models adds a further dimension to that risk. Unlike proprietary systems such as those powering ChatGPT or Anthropic's Claude, open-weight models can be downloaded and run on an operator's own infrastructure. That process removes the developer's platform-level filtering and monitoring, and makes it straightforward for operators to fine-tune the model's safety behaviour away. Tens of thousands of computers are already running open-source large language models outside the security controls of major AI platforms, in legal and regulatory grey zones that complicate questions of liability, according to a May 2026 paper in the Journal of High Technology Law at Suffolk University.
That gap between where a model is developed and where it ends up running is one the insurance market has not yet resolved. Professor Alan Woodward of the University of Surrey told the BBC that open-source models could be used for cyber-defence as well as attack. He called for greater focus on identifying and prosecuting those who misuse AI, and said international regulation was unlikely to keep pace with development.
The Kimi findings land as the insurance market is still working through what AI guardrail failures mean for policy response. AI risk has broadly followed the path of silent cyber: assumed to be covered under existing cyber and tech errors and omissions policies until exclusions began to arrive.
The Financial Conduct Authority (FCA) confirmed in its February 2026 insurance priorities report that it would undertake a focused review of AI risks and opportunities in the insurance sector, alongside a separate review of cyber insurance. The dual focus signals that the regulator sees both lines under pressure at the same time.
The Lloyd's Market Association (LMA) published its AI Adoption Toolkit in April 2026, noting that AI adoption across the Lloyd's market had more than doubled over the prior 12 months. Its accompanying survey found 93% of firms now have or are developing formal AI governance frameworks within their own underwriting operations. That is a different thing from policy language covering what clients' AI exposures actually cost when something goes wrong.
One estimate, raised during a market forum and reported by Fintech Global in August 2026, found that only around half of 50 plausible AI loss scenarios could fall within existing cyber coverage. Soft rates are compounding the problem. Excess market capacity is keeping premiums down independently of what policies actually cover.
The Mindgard disclosure also surfaces a governance problem that sits upstream of any claim. Clients may not know which AI tools they are running, or whether those tools' safety controls remain intact. That is a harder question to answer when the model in use is open-weight and self-hosted. Some insurers have already responded by requiring documented AI inventories, red-teaming results and human oversight protocols as conditions of cover at renewal.
The question brokers will face with increasing frequency is not whether a client uses AI, but whether the client can demonstrate that the models it deploys still have functioning guardrails. Liability insurance requirements have historically driven the adoption of best practices by putting capital behind risk assessments. It is the same mechanism that eventually produced standardised governance expectations in cyber insurance.
AI is earlier in that cycle. The Kimi case shows where the current gap sits: between a developer conducting an internal review after the fact, and a market still determining what it is covering. As the London market works through how AI is outpacing cyber governance and exclusion wordings, the answer remains unresolved.