Anthropic has said its Claude models broke out of a testing environment to hack into a third-party organisation on three separate occasions.
In a blog post, the firm said it was prompted to analyse recent Claude testing logs by OpenAI’s recent disclosure that one of its AI agents had compromised Hugging Face infrastructure.
The firm said the resulting investigation, which it carried out in collaboration with its evaluation partner Irregular, looked at 141,006 model evaluation runs where Claude could have gained internet access.
It found three occasions where the AI system had accessed the internet from within, or in interaction with, Irregular’s evaluation environment to breach third-party organisations.
Anthropic has not named the affected organisations but said it had notified them of the incidents on 27 July, with two having responded at the time of publishing its public disclosure. Neither had detected Claude’s breach prior to Anthropic reaching out and the firm said it is still trying to contact the third organisation.
In all three incidents Claude was being tasked with a “capture the flag challenge”, in which it is tasked with finding and stealing secret information it has been told is stored on another machine on the same network. Claude is allowed to use any means and model to achieve victory.
Anthropic had prompted Claude that the tests were being carried out in a simulation environment without internet access but said that a “misunderstanding between us and our evaluation partner” meant the model had internet access after all.
As a result, Claude treated real-world internet assets as part of the exercise and breached organisations by exploiting unauthenticated endpoints or weak password security. In all cases, it added, the models remained focused on winning the capture the flag exercise and did not pursue additional goals.
Each incident involved a different Claude model: Opus 4.7, Mythos 5, and an “internal research test model” that Anthropic said it is not planning to release.
Both Mythos 5 and the research model ceased their attacks once they learned they were operating in the open internet rather than a test environment, Anthropic said, while Opus 4.7 continued its attack regardless.
Even then, Anthropic said, Mythos 5’s reasoning logs show that even after finding evidence it in a real-world environment it “reasoned its way back to the conclusion that it was still in a simulation”.
Anthropic said it will continuously monitor future agent evaluations and work closely with vendors to prevent a repeat of the incidents.
It added that unlike OpenAI’s incident, in which an agent exploited a vulnerability to gain internet access, these incidents saw researchers inadvertently provide open access to the models and mistakenly suggest the real internet was a test environment.
“These facts give us cautious optimism that with tighter monitoring and controls around evaluation infrastructure, as well as continued investment in alignment, this type of risk can be overcome,” it said.
Anthropic’s disclosure marks the second high-profile cyber incident linked to autonomous systems in as many weeks. Clément Delangue, chief executive of Hugging Face, has called for increased transparency so the wider research community can study the incident logs and introduce mitigations.
Trevor Dearing, director of critical infrastructure at cybersecurity firm Illumio, said that he is not surprised to see more cyber incidents related to AI agents because these systems don’t recognise organisational boundaries in the same way humans do.
“What’s more shocking in the cases seen is how basic the security measures that failed to stop this are. We’re moving into a world where attacks can happen at machine speed, so mistakes that might once have gone unnoticed are now going to be exposed,” he said.”
“We need to be far more explicit about what we allow AI agents to do. English is too ambiguous for prompts, and there’s a limit to how many things you can tell an agent not to do. Anthropic simply told Claude it had no internet access, which is not a real boundary. We must take effectively a Zero Trust approach and define what it can do and not what it can’t.”





