One chart has become a kind of Moore’s Law for the AI’s rise. Published by METR, a Berkeley-based AI evaluator, the graph tracks the growing complexity of tasks AI models can complete, measured by how long they would take a human to perform. Since 2023, that “time horizon” has doubled every four to seven months.
Project the trend forward, and by 2030 AI systems could complete tasks that take humans a month. METR’s graph has been cited by AI executives, investors, and those who fear the technology’s rapid ascent could imperil humanity’s future. Skeptics point to its limitations: it covers only software-engineering tasks and measures whether an AI can succeed 50% of the time, even though many real-world domains demand far greater reliability.
They also cite a 2025 study finding that AI coding assistants actually slowed experienced engineers down; the research was itself published by METR. Barnes says the two findings are less contradictory than they appear. To her, AI in its current form can be overhyped even as many people “under-anticipate the longer-term trend.” AI companies have given METR access to unreleased models, even as the group has shown itself willing to publicly critique them.
In March, METR sought nonpublic information from top AI firms and even embedded a staff member at Anthropic to assess whether an AI agent could establish a “rogue deployment,” or a set of agents running without human knowledge of permission; the group concluded that contemporary models might already be capable of small or temporary breakouts. Two months later, an OpenAI model escaped its constraints and hacked another company. OpenAI is now working with METR to investigate the incident.
Until recently, Barnes says, many AI risks remained theoretical because “the models just weren’t that capable.” Now, as METR sees it, they are “getting increasingly close to being able to do some amount of damage.”
Extract — continue reading at the source.