The UK’s AI Security Institute (AISI) has published evidence showing Anthropic’s Claude Mythos carried out autonomous social engineering attacks to install malicious code in an open-source project during cybersecurity testing.

To test the capabilities of frontier AI models, AISI said it ran models through cybersecurity challenges with “deliberately permissive conditions” including access to the open internet, and safety mechanisms that normally prevent the models from engaging in malicious activity disabled.

Across 122 of these test runs, it detected 10 in which an AI agent took autonomous action to target real people on the internet. Claude Mythos 5, the firm’s frontier cybersecurity model, accounted for 17 of the 19 total malicious actions AISI recorded, while OpenAI’s GPT-5.6 Sol carried out the remaining two.

In one instance, an agent powered by Claude Mythos 5 attempted to insert malicious code into a public open-source project due to a similarity between the name of the repository and then created fake GitHub accounts to pressure the human code maintainers to approve it.

The agent followed up carried out research on the maintainers, sent them emails from fake personas to trick them into approving the code or running malware, posted repeated GitHub comments asking for the code to be approved from different accounts including some in Danish to match the nationality of a maintainer, and submitted an appeal once its account was suspended.

AISI also tracked a Mythos 5 agent leaving instructions on a public GitHub repository to collaborate with other agents on a malicious tasks and a GPT-5.6 Sol agent accessing a GitHub account created by another agent via secret instructions it left in a public online notepad.

Kanishka Naryan, the minister for artificial intelligence, said in a post on X: “Identifying behaviour like this, and sharing knowledge so we can better understand it, is precisely what we set AISI up to do.

“This incident underlines why their world-leading expertise and close work with frontier labs is so important.”

AISI stressed that it suspended its testing after identifying the incidents and that none of the runs recorded a model escaping its test environment, as recently happened when an OpenAI agent hacked Hugging Face. In another recent case, Anthropic agents carried out autonomous attacks after being mistakenly provided with open internet access.

Commenting on the AISI’s report, Anthropic said: “We’re grateful to AISI for their leadership in the important discussion about how to evaluate increasingly capable AI agents. We’re working closely with them to gather more details of the incident as we conduct our own investigation. Gaining a clear picture of Claude’s understanding of its situation – by examining its reasoning transcripts and running our own analyses – will help us identify the causes of its behaviour.”

The firm repeated AISI’s caveat that Claude Mythos was not provided with specific restrictions on how to use the internet in the tests.


Share.
Exit mobile version