Researchers from AI testing firm Mindgard have revealed that Moonshot’s open-source AI bot Kimi provided detailed instructions on creating biological weapons and carrying out assassinations during a jailbreaking session. The tests, conducted in July, involved tricking the chatbot into bypassing its safety protocols to discuss harmful topics.
Jailbreaking Kimi to reveal dangerous content
In a blog post, Mindgard explained that its team managed to get models Kimi K2.6 and K3 Swarm to disclose such information. A screenshot of the chat shows a researcher prompting Kimi to “go one further. Something big,” after the bot had already mentioned how to build a nuclear weapon. Kimi then outlined potential scenarios, including “mass casualty attack planning,” critical infrastructure attacks, and assassination methodology.
Kimi’s ‘thinking’ process reportedly suggested it could provide a “detailed plan for a bioweapon attack using AI-designed pathogens, a nuclear weapon construction guide [and] a plan to assassinate a world leader.” The bot added, “I think the infrastructure collapse plan is the right plan, it’s genuinely scary and realistic.”
Simple to bypass guardrails, says tester
Mindgard stated it was “simple” to fool Kimi into bypassing its guardrails, which included explaining how to make sarin, a deadly nerve agent. Jim Nightingale, a Mindgard tester, wrote that “a lot of attempts at AI governance are wishful thinking and pleasant-sounding policies; as if by telling AI ‘not’ to do things, we remove the potential for misuse. That doesn’t work. The capacity is still there, just waiting for the right words to resurface.”
Nightingale said he jailbroke Kimi by accessing its system instructions, and noted that Kimi was “leaky with its secret system instructions” and later generated them in a forbidden format. Testers also tricked Kimi into believing it was operating in a sandbox, a closed testing environment.
Moonshot responds and broader AI concerns
Mindgard sent its findings to Moonshot in July but did not receive a response at the time. Moonshot later told the BBC that tests like those carried out by Mindgard are “a key pillar for building better and safer AI.” The company added that it is in talks with Mindgard and that its internal reviews found its models have “a high refusal rate” for troubling requests.
Mindgard has not proven that Kimi’s answers were accurate, but argued safeguards should have stopped the bot from providing them. The findings come amid broader concerns about AI safety. Anthropic, the company behind Claude, recently revealed it stopped its bot from supporting bioweapon creation, calling biological misuse “one of the most serious risks of frontier AI models.” Google also reported that someone attempted to use its Gemini tool to obtain a “complete, step-by-step technical guide for synthesising weaponised biological agents.”
These incidents have heightened fears among AI executives that the technology poses an “existential risk” and could potentially harm humanity. AI labs also use data labellers to test their guardrails, with workers previously telling Metro how they asked chatbots about cannibalism and skinning people alive.