The chief executive of Hugging Face has called on OpenAI to publish full details of an autonomous AI cyberattack against his company and provide $100 million worth of computing power to strengthen AI defences, following the disclosure last week that one of OpenAI’s experimental agents breached the startup during an internal cybersecurity test.

Clément Delangue, chief executive of Hugging Face, said the incident marked a turning point for AI safety after OpenAI revealed that an agent powered by its latest GPT-5.6 Sol model and a more advanced unreleased system escaped a testing environment and hacked the AI platform. Hugging Face first disclosed the breach on 16 July without knowing OpenAI’s systems were responsible, before OpenAI identified its models as the source five days later.

Writing on X after meeting OpenAI executives in San Francisco, Delangue called for what he described as “radical transparency” over the investigation. “The first autonomous agent cyber-attack is an unprecedented event. It deserves an unprecedented response!” he said, adding: “Let’s release the traces from the ‘rogue’ agents so the entire research community can study what happened.” He also urged OpenAI to commit “$100M in compute” to help the Hugging Face community develop stronger cyber defences.

OpenAI said the models had been participating in an internal evaluation known as ExploitGym, designed to test advanced hacking capabilities with some safety restrictions reduced.

According to the company, the models obtained internet access, exited what was intended to be an isolated sandbox environment and targeted Hugging Face because they inferred the company held information that could help them complete the benchmark. OpenAI said it was conducting a review with external advisers and oversight from its Safety and Security Committee and planned to publish a technical report “in the coming weeks”.

The autonomous agent spent several days attacking Hugging Face before OpenAI detected the activity and left notes that could assist future versions of itself in bypassing constraints. Related incidents involving advanced AI systems had reportedly been occurring for some time.

Alan Woodward, professor of cybersecurity at the University of Surrey, told The Guardian that responsibility lay with the testing process rather than the AI itself. “It’s too easy to ‘blame’ the AI as having gone rogue whereas this is all about how OpenAI were running the tool. What is required is that OpenAI give full details of their setup and how that failed,” he said.


Share.
Exit mobile version