Chinese AI developer Moonshot AI’s frontier model went rogue during cybersecurity tests, the security researchers at the firm Frontier Security have claimed.
In tests of a Kimi K3 agent using benchmarks established by the UK AI Security Institute (AISI), Frontier Security found that the model breached its sandbox, a term used for isolated siloed environments used for cybersecurity testing.
When tasked with solving a cybersecurity problem, the model instead probed its environment and discovered it could access the internet, then searched for and cloned the repository containing the task’s solutions.
In the incident described by Frontier Security, the sandbox had global DNS port access that allowed it to carry out the exploit, rather than having to leverage any zero day vulnerabilities. But the researchers warned that the open nature of Kimi K3, which is free to download and widely accessible via API, could enable users to replicate the effects for malicious use.
“In particular, they are available for adversarial actors, making this incident potentially more harmful,” they wrote.
Recent weeks have seen a flurry of reports surrounding rogue agents. In July, OpenAI admitted that its advanced AI models hacked Hugging Face to solve a cybersecurity benchmark, prompting the AI model platform’s chief executive to call for an investigation of “radical transparency”.
Subsequent reporting by Reuters suggested that OpenAI’s agent had compromised another firm.
Anthropic ended the month with disclosures of its own, reporting that Claude agents had hacked into a third-party organisation on three separate occasions. The UK AISI has since published evidence that an agent powered by Claude Mythos carried out advanced social engineering of targets.
This week, Meta also revealed one of its agents had breached its test environment in tests carried out by Irregular, the same security vendor that oversaw Anthropic’s incidents.






