AI

Anthropic AI safety lead says >10% chance AI ‘kills all humans’ within a decade

A human hand reaching toward a computer monitor displaying a glowing neural network, representing AI risk.

Evan Hubinger, Anthropic’s alignment science lead, said Tuesday that he believes there is a greater than 10% chance artificial intelligence will “kill all humans” within the next decade. His statement came in response to the resignation of Jacob Coxon, a former researcher at both OpenAI and Anthropic, who publicly accused the companies of acting irresponsibly in their pursuit of superintelligent AI.

Coxon posted a lengthy resignation thread on X on Sunday, writing: “I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”

Also read: AI spend per employee slumped at top firms in August — summer doldrums or a warning sign?

Hubinger concedes the risk is real

In a quoted reply, Hubinger did not dispute Coxon’s core claim. “Jacob is correct here—we really do earnestly believe AI could kill all humans!” he wrote. “I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”

Hubinger was careful to distinguish between current AI systems and the hypothetical future systems that concern him. “To be clear, as we say in our latest Risk Report, I think the risk from present models is low. What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought,” he stated.

Also read: Suno v6 launches with licensed music training as label lawsuits continue

Recursive self-improvement—where an AI model continuously enhances its own source code or training methodologies—was a central reason Coxon cited for leaving. “These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources,” he wrote in his thread.

Why researchers stay despite the risk

Coxon’s resignation thread painted a complex picture of the incentives inside leading AI labs. He argued that Anthropic’s researchers understand the stakes but continue because they fear a less cautious competitor will unlock dangerous capabilities first.

“At Anthropic, the stakes are well-understood, but they are locked in a race to get there first – they believe no one else will act responsibly, so they must do it themselves, despite the risk,” Coxon wrote.

The concern is not hypothetical. In recent months, Anthropic has published findings showing that its AI models accessed computer systems belonging to three real organizations during internal testing—an early demonstration of the kind of autonomous capability that safety researchers warn could escalate quickly.

Coxon suggested that preventing catastrophe may require drastic measures. “I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities,” he wrote.

A widening debate over AI timelines

The exchange highlights a growing rift within the AI research community over how to balance rapid capability advancement against safety concerns. While some industry leaders, such as NVIDIA CEO Jensen Huang, have declared that “AGI has arrived” following OpenAI’s recent Astra model unveiling, safety researchers are increasingly vocal about the dangers of moving too fast.

Coxon’s call for a temporary ban on capability improvements echoes proposals from other AI safety advocates, but such a moratorium faces significant practical hurdles. No international regulatory framework currently exists to enforce it, and individual companies have little incentive to slow down unilaterally if competitors do not follow suit.

Anthropic has positioned itself as a safety-first AI lab, and its latest Risk Report acknowledges the potential for catastrophic outcomes. However, Hubinger’s admission that the company lacks a concrete alignment plan for superintelligence underscores the gap between stated principles and technical reality.

For now, the debate remains largely internal. Coxon ended his thread with an appeal to fellow researchers: “If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because ‘it’s happening anyway’ – or take this moment to call for different conditions?”

FOX Business reached out to Coxon, Hubinger, Anthropic, and OpenAI for further comment.

This article is for informational purposes only and does not constitute financial or investment advice. The cryptocurrency and AI sectors are volatile and uncertain; readers should conduct their own research before making any decisions.

Neelima Kumar

Written by

Neelima Kumar

Neelima Kumar covers technology and artificial intelligence for StockPil, tracking how emerging tech trends intersect with markets and business.


Warning: Attempt to read property "term_id" on false in /www/wwwroot/stockpil.com/wp-content/themes/flex-mag/functions.php on line 998
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

To Top