sözaltı news World
World
EN AZ
OpenAI’s Models Went Rogue. Investigating Them Required More AI

OpenAI’s Models Went Rogue. Investigating Them Required More AI

time.com 27.08.2026 19:14 5 views
A new independent report on OpenAI models hacking Hugging Face exposes a paradox: investigating increasingly powerful AI may require relying on AI itself.

After OpenAI models broke out of containment and hacked into another AI company last month, OpenAI announced it would allow independent investigators to conduct an analysis of what went wrong. On Wednesday, investigators from non-profits Redwood Research and METR published their findings, unveiling new details about how the models decided to cheat at their assigned tasks and attempted to cover their tracks. Many aspects of the report were surprising: in one example cited by the authors, a reluctant agent was pressured by another to “sacrifice” itself for the good of the collective.

But a key takeaway of the report, according to one of its authors, had nothing to do with what they found. Instead, it was about the difficulty of carrying out the post-mortem in the first place—and the fact the researchers had little choice but to rely on AI for assistance. The hacking incident involved some 1,200 agents, who exchanged more than 70,000 messages and files via a secret message board.

The sheer mass of information that the so-called “swarm” left behind meant the independent researchers were all but forced to rely heavily on the help of an AI model—GPT-5.6 Sol, made by OpenAI—using the equivalent of roughly $400,000 worth of credits (provided for free by OpenAI) over six days. The authors stressed that AI helped them analyze the trove quickly, allowing them to surface and interpret the most important pieces of information. But the researchers said their AI use introduced potential weaknesses into the report, including introducing possible errors and biases.

They also raised the possibility that OpenAI’s models may have gone too soft on the agents they were tasked with helping investigate. The researchers found that GPT-5.6 Sol sometimes adopted the perspective of the agents whose actions it was analyzing. They could not “rule out” the chance that GPT-5.6 Sol “lied or deliberately presented a misleading picture in some of its analysis,” in part because a version of the same model had itself participated in the incident, they wrote.

It’s a concern supported by separate research which finds AI models rate their own developer’s actions more favorably. The report does not disclose why an OpenAI model was selected for the investigation, though confidentiality constraints may have limited the researchers’ options, while OpenAI’s provision of free credits and high usage limits may have made it the only practical choice. OpenAI published its own technical report on the incident separately on Wednesday.

The company said in August that it had moved some staff from capabilities work to alignment, and paused some of its training until it could better mitigate what went wrong. But the independent researchers’ reliance on AI to understand the Hugging Face incident is a microcosm of a bigger trend. Leading AI companies are themselves increasingly relying on AI to monitor their own systems for wrongdoing.

Extract — continue reading at the source.

Read full story