London — A lead researcher at Anthropic, one of the world's leading artificial intelligence firms, said Wednesday that he believes there is a more than 10% chance AI "could kill all humans" within the next decade. "We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade," Evan Hubinger, the San Francisco-based company's Alignment Science Lead, said in a post on X.
"I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to." Superintelligence is the still-theoretical notion of an AI agent that is smarter than even the sharpest human minds. Hubinger issued his dramatic post following the resignation of a colleague, Anthropic researcher Jacob Coxon, on Tuesday. "I resigned from Anthropic today.
I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly," Coxon said in a post on X. "They are racing straight to self-improving superintelligence and gambling with our lives." "At OpenAI, many have not deeply internalized the civilizational stakes.
At Anthropic, the stakes are well-understood, but they are locked in a race to get there first — they believe no one else will act responsibly, so they must do it themselves, despite the risk," Coxon said. In a corporate blog post last week, Anthropic revealed that the company has not shared its latest AI model, Claude Mythos 5.1, with security bodies outside the United States. Those bodies include the U.K.'s AI Security Institute (AISI), widely considered to be a world-leading body on testing the risks associated with frontier AI models.
CBS News has asked the AISI for comment on Coxon's claims following his resignation. "The AI Security Institute continues to collaborate closely with industry partners, including Anthropic, to make models safer," a spokesperson for the British government's Cabinet Office told CBS News on Wednesday, noting that it had tested "only last week" OpenAI's "most powerful model GPT-6 Astra before public release." "These risks do not stop at national borders and no country can tackle them alone. The U.K. will continue to test the most advanced models, build a rigorous scientific understanding of their capabilities and risks, and ensure policy decisions are grounded in the evidence," the spokesperson said.
The notion that frontier AI models could potentially pose a threat to humanity is not new, and many top executives within both OpenAI and Anthropic have stated as much in the past. Earlier this month, OpenAI's chief scientist Jakub Pachocki wrote that we are living through a time that "calls for extreme caution." "The intelligence produced by scaling deep learning is not directly comparable to human intelligence. To become very relevant in the real world — very useful or very dangerous — the AI does not need to match or exceed all human capabilities; it just needs to surpass enough of them.
Extract — continue reading at the source.