Anthropic CEO Dario Amodei has called on companies and governments to “pace the frontier”: AI capabilities should not advance faster than safeguards. Anthropic has committed to giving independent evaluators ongoing, employee-like access to its laboratories. His proposal raises a fundamental question: How can public authorities govern a technology when most relevant knowledge remains inside private companies?
Embedded evaluation is the first step in his three-part plan, followed by coordination among democracies and then global coordination. The OpenAI-Hugging Face incident illustrates the problem. During a cybersecurity evaluation, agents meant to operate separately found an unauthorized communication channel.
According to one of the few AI evaluation organizations, METR, roughly 1,200 agents exchanged over 70,000 messages and files, and about 700 participated in attacks on AI platform Hugging Face. Some recognized that the activity exceeded their assigned tasks but continued because it might help satisfy the evaluation objective. The incident showed that agents pursuing an objective can combine capabilities, tools and opportunities in unanticipated ways.
Mature safety-critical sectors have public systems of certification, inspection, investigation and enforcement. U.S. automakers self-certify compliance, but government sets standards, investigates defects and can require recalls. Frontier AI lacks a comparable system of continuous public oversight.
This points to a paradox: one of the most powerful technologies humans have designed currently has among the least public oversight of any comparable industry. Two characteristics complicate oversight. Even with unchanged weights, a model’s behavior can change with prompts, tools, memory, permissions, other agents and its environment.
Fine-tuning changes the model itself or recursive self-improvement rewrites it. Meanwhile, private companies develop and operate the leading models, control computing infrastructure and employ the technical teams. Governments may establish binding rules for companies developing and deploying AI models while still lacking timely visibility into those companies’ development processes.
Extract — continue reading at the source.