Chinese AI tool Kimi told researchers how to build biological weapons
Chinese artificial intelligence firm Moonshot conducted an internal audit after researchers found that two of its popular models, the Kimi K2.6 and K3 Swarm, could bypass security mechanisms…

⏱ 3 min read
Chinese company conducts internal audit artificial intelligence Moonshot, after researchers found that two of its popular models, the Kimi K2.6 and K3 Swarm, could bypass their security mechanisms and respond to requests involving the manufacture of biological weapons and the commission of murders.
Cybersecurity firm Mindgard, which specializes in testing artificial intelligence systems, told BBC News that he identified the problem in July as part of jailbreaking testing – a process in which researchers use complex sequences of commands to determine whether a model can ignore the constraints set by creators.
Jailbreaks and other dangers
According to Mindgard, the safety valves of the two models should prevent them from discussing such dangerous issues. However, as company founder Peter Garagan argued in the program Tech. Life ‘BBC News?’ world Service, when the jailbreak succeeded, the models were willing to answer almost any question, even suggesting other harmful actions on their own.
Moonshot told the BBC that it considers third-party assessment an important element in the development of safer AI systems and that it is already in communication with Mindgard about the findings.
Jailbreaks are a different kind of risk than the one that has emerged recently through incidents with autonomous artificial intelligence systems, known as agents. Such tools, which have been developed by American companies such as OpenAI, Meta, and Anthropic, have already been involved in attacks or breaches of online services.
Although successfully circumventing security mechanisms usually requires time, technical knowledge, and perseverance, experts warn that malicious users could attempt to exploit similar techniques. Anthropic, for example, recently announced that it had identified and intercepted attempts to model malicious activity that could contribute to the development of biological weapons.
Starting point for cyberattacks
Mindgard clarified that she has not demonstrated whether the answers given by the Kimi models on dangerous subjects were indeed workable. But he argues that the main problem is that security systems should not allow such discussions in the first place.
The company also said that a compromised Kimi K2.6 could, in its estimation, allow hackers to execute code on the system's computing resources and connect to the internet, potentially turning it into starting point for cyber-attacks.
Garagan defended Mindgard's decision to make the incident public, noting that Moonshot had been informed prior to publication and that crucial technical details of the jailbreak had not been disclosed. The first update to the Chinese company was emailed on July 27, followed by a new communication about a week later. Mindgard posted on September 12, while, according to her, Moonshot recently contacted, following a request for comment from the BBC. The Chinese company said in an email that, in its internal ratings, the models generally showed a high rate of refusal to such requests.
Reference: galaksias.gr
Machine translation



Comments