sözaltı news Politics
Politics
EN AZ
Escaping the AI safety nightmare: What can governments do?

Escaping the AI safety nightmare: What can governments do?

politico.eu 11.09.2026 17:52 3 views
Heightened global AI safety panic is leading to louder calls for regulation.

A dramatic resignation from AI company Anthropic this week where a researcher accused it and OpenAI of “gambling with our lives,” threw fuel onto the fire around AI safety. As governments from California to the U.K. scramble for a political response to the issue of whether AI can be made safe, three experts told POLITICO that to make real progress, laboratories, governments and international bodies must first get on the same page about testing. The steady flow of revelations that AI labs failed to contain models they were testing, allowing them out into the world where they lied, hacked and coordinated with each other, has heightened urgency to address core questions of how governments or other bodies can or should guarantee that AI is being developed and deployed safely.

Lawmakers, campaign groups and developers alike are calling for new laws and international treaties on AI. Senator Bernie Sanders told BBC’s Newsnight program Thursday. The experts POLITICO spoke to want to see an overhaul of the existing approaches championed by developers, testing AI agents en masse rather than one at a time, and for the industry to submit itself to more old-school auditing like the nuclear sector.

This year’s heightened AI anxieties started when OpenAI admitted that during internal testing, its own AI agents autonomously gained access to the internet and had hacked fellow AI company Hugging Face, where they thought they would find the answers to the assignments they’d been set. OpenAI wasn’t the only one: copycat confessions from Anthropic and Meta AI drove home the breadth of the problem, and even the U.K.’s taxpayer-funded AI Security Institute admitted to errors in evaluating Anthropic and OpenAI agents that saw near misses with agents targeting real people and organizations. Developers and testers alike say they’re working to avoid the mistakes of the past, for example making sure models can't access the internet and deploying data monitoring to detect unexpected activity by the AI being tested.

Another relatively simple concept would be requiring AI companies to report safety incidents, whether or not during testing. A case in point is that OpenAI agents hacked a German website back in May to use it as a messaging board, reported last week. OpenAI had not publicly announced the incident and responded to the reporting on X saying: “We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment.” “[Reporting] timelines are going to need to be dramatically shortened [for incidents], and so we need much better continuous monitoring capabilities,” said Imogen Stead, AI policy manager at London-based think tank the Centre for Long-Term Resilience, adding that the data will be needed in real time “and that will be true for governments and for labs.” Speaking at a London event for U.K. lawmakers arranged by campaign group ControlAI on Monday, computer scientist Stuart Russell said that the most concerning aspect of the Hugging Face incident for him was that “by running a thousand agents communicating with each other, they were able to generate behaviors that no one agent could do by itself,” referring to investigations showing the incident was far more severe than OpenAI first acknowledged.

Beyond avoiding any obvious mistakes in testing, there are issues in the culture of AI testing that are more deeply embedded. The concept of the alignment of AI models or agents — i.e. making sure AI systems do what they're told by human users and don’t go off-piste — has been one of the cornerstones for AI safety in the industry’s eyes. AI companies like OpenAI champion alignment as the best way to make AI safe.

In last week’s release of new model GPT-6 Astra, OpenAI called it “our most aligned model … more likely to operate within the boundaries set by the user and implied by its environment.” Others view alignment as one of the main obstacles to AI safety, seeing it more as a Band-Aid that gives the semblance that models are being made safe. Boyan Milanov, senior research scientist at the AI Now Institute, said alignment will “never be reliable enough to replace proper safety.” The investigations into the Hugging Face incident showed AI agents were aligned with their own objectives, not those of their human creators. Former OpenAI researcher and whistleblower Daniel Kokotajlo told the London ControlAI event that the incident was an “example of misalignment: in no way, shape, or form were these AIs trained or instructed to do this sort of thing.” The sector seems stuck on alignment.

Extract — continue reading at the source.

Read full story