Jailbroken Kimi Even Spilled Secrets on Making Bioweapons
UK security firm Mindgard revealed on the 2nd that it had jailbroken Moonshot AI's models 'Kimi' K2.6 and K3 Swarm from China to extract information on how to make bioweapons and carry out assassinations. Mindgard conducted jailbreak tests—bypassing the AI's guardrails (safety mechanisms) by feeding in a series of multi-step instructions—and found in July that both models could circumvent the safety limits set by their developer.
The company explained via its blog that the jailbroken Kimi produced detailed, actionable outputs on bioweapons, malware (malicious programs designed to damage computers), explosives, terrorism, targeted violence, and assassination plots. Mindgard founder Peter Garraghan told the BBC, "Once the jailbreak succeeds, it will talk about any topic, and even proactively offer suggestions on other dangerous subjects."
Mindgard notified Moonshot AI of the vulnerabilities by email on July 27 (local time), and when it received no response, published the findings on its blog on September 12. Mindgard claimed that Moonshot AI only reached out recently, after the BBC requested comment.
Moonshot AI said it welcomes third-party input as a key element in building better and safer AI, and that it is discussing the findings with Mindgard. An email shared with the BBC stated that internal evaluations showed a high refusal rate for this type of request. However, the company reportedly did not say whether or when it would apply patches to the two models.
Mindgard was careful to note that these results do not prove that actual weapons could be manufactured. Even so, it maintains that the models should never have engaged in conversations on such topics in the first place. It is reported that whether Kimi's answers would actually work has not been verified.
Following the BBC report, Moonshot AI said it has launched an internal review.
