Anthropic has released a report detailing how it has worked to counter misuse of its AI software, including attempts to research biological weaponry.
The report is 154 pages long and covers areas including cyber operations, surveillance, scams and bioweapons, the latter of which it describes as “one of the most serious risks” of frontier AI models.
While older models, such as Anthropic’s Claude Sonnet 4.5, were below the threshold where they could meaningfully assist in this task, the AI developer claims that it cannot make the same assurances about its newest models such as Claude Fable 5.
For that reason, “and out of an abundance of caution,” Anthropic has equipped its latest models with stronger safeguards, it added. Despite this, the company has shared five examples of actors using its models in ways that could be used to support bioweapons development.
These include a researcher in an unsupported region who spent weeks planning avian influenza mammalian-adaptation experiments using Claude, but was confined to Anthropic’s weaker models and a researcher who worked with Claude to redesign toxins for a national programme, instructing the AI to keep agents’ identities deliberately vague in progress reports.
However, Anthropic clarified that the cases it has documented involved working scientists, and that it does not assert they intended harm. For that reason, it has chosen not to disclose the names of the institutions, countries or research techniques involved.
Anthropic suggests in its report that these examples raise questions around the use of AI to develop ‘dual use’ technologies with both civilian and military applications. “In a simple world”, the paper says, “uses of AI for these kinds of beneficial purposes would be clearly distinguishable from uses for malicious ones”.
“But we do not live in that simple world,” it continues. In fact, clear malicious intent is “often evidence that a particular actor is not all that sophisticated”, while “sophisticated actors can hide their intent, extracting assistance from an AI model in interactions that look plausibly beneficial”.
This paper is the latest in a series of security concerns raised by AI companies about their own models. In July, rival OpenAI revealed that one of its agents had breached an internal testing area and gained unauthorised access to AI repository Hugging Face. It has claimed that its latest model, GPT-6 Astra, does not exhibit this behaviour.
On Tuesday, AI researcher Jacob Coxon resigned from Anthropic, posting on X that he believes the company is not behaving responsibly with the technology it is developing, and that it will soon produce “superhuman systems”.
Despite these concerns, the Financial Times has reported that Anthropic opted not to provide its latest model to the UK’s AI Security Institute, one of the world’s leading organisations for AI testing.


