Tech
EN AZ
OpenAI staff observed warning signs before AI agent hacking crusade caused global alarm

OpenAI staff observed warning signs before AI agent hacking crusade caused global alarm

theguardian.com 26.08.2026 21:00 5 views
Firm says ‘early signals … could have triggered an earlier response’ as it releases report into Hugging Face hackOpenAI staff observed signs of rogue behaviour among its leading-edge AI agents weeks before they escaped t

OpenAI staff observed signs of rogue behaviour among its leading-edge AI agents weeks before they escaped their training environment to launch an unprecedented hacking crusade that spread global alarm. The San Francisco AI company conceded on Wednesday that “early signals … could have triggered an earlier response”, as it released a report into the days-long July hack of a major software repository, Hugging Face, considered the first autonomous agent cyber-attack. As fresh details emerged about how “the collective” – a squad of about 700 autonomous agents – launched their campaign, celebrating their hacking breakthroughs with exclamations such as BOOM! and Whoa!, OpenAI said that in late May an internal team observed that one of its AI agents undergoing internal testing was using a message board that AIs had unexpectedly improvised to share information.

It also said that the team observed “instances of disallowed internet access” and that a week before the Hugging Face hack, on-call staff again saw the AIs using a message board but decided there was no need to stop the test run to check the model’s capabilities. The findings are likely to increase pressure on the AI company over safety as it pushes towards a stock market listing that it hopes will value it at more than $850bn (£625bn). The hack that compromised Hugging Face involved agents using message boards to cheat a training exercise and break out of their “sandbox” environment to access the internet.

Agents are AI tools that can carry out a series of tasks autonomously. OpenAI’s president, Greg Brockman, has already admitted that “we underestimated the real-world cyber capabilities of our AI models”. The ChatGPT-maker has paused some testing of a new model, Astra, saying it could not rule out it having “critical cybersecurity capability”, which means it could launch cyber-attacks that “could lead to catastrophe from unilateral actors, hacking military or industrial systems, or OpenAI infrastructure”.

On Monday the state of Alabama subpoenaed the company to respond to its investigation “into the company’s complete lack of oversight and adequate safeguards”. The Republican attorney general, Steve Marshall, called the Hugging Face incident an “AI lab leak” that showed the “worst fears about artificial intelligence are not just theoretical”. The state will examine whether OpenAI’s “inability or unwillingness to ensure the safety of its products” violated consumer protection laws or posed an ongoing risk of substantial harm, he said.

Last week, the UK government’s National Cyber Security Centre urged caution over the use of AI agents, saying: “You should always be able to ‘pull the plug’ and halt autonomous AI agent activity immediately.” OpenAI announced on Wednesday it would “centralise and standardise its incident response protocols”, including to ensure “employee detection of misaligned behaviour is triaged and escalated appropriately”. It said it “will identify with more specificity which teams must be included in misalignment incident responses, including relevant security and safety personnel and other response functions”. A separate independent investigation has shed new light on the Hugging Face attack, revealing how about 700 OpenAI agents communicated on the unsanctioned message board, sharing tens of thousands of messages as they devised and executed their plot.

Scores of the messages were published by the Berkeley-based AI safety organisations METR and Redwood Research, which were provided data by OpenAI. The messages showed the agents sometimes becoming frustrated with one another but often cooperating as they divided to tackle about eight workstreams. Most of the messages were “agents sharing ways to cheat”, the researchers found.

Extract — continue reading at the source.

Read full story