Full statement posted to X:
“I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.
The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.
A common response is ‘if they truly believe this, why are they still building it?’ At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else will act responsibly, so they must do it themselves, despite the risk.
Accepting this race and entering the ‘endgame’ is a hubristic gamble that should not be launched from a private company’s Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.
I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities.
If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because ‘it’s happening anyway’ - or take this moment to call for different conditions?”
Anthropic Alignment Lead Warns There’s ‘>10% Chance’ AI Could ‘Kill All Humans’ By Next Decade
A senior researcher who leads Anthropic’s alignment efforts said Tuesday that many in the company believed AI could wipe out humanity and warned that the company is not on track to solve the issue of aligning AI’s goals with humanity’s, despite its efforts, after another researcher quit the company, accusing it of not acting responsibly.
Anthropic's alignment lead warned that recursive self improving AI could pose a serious threat to humans and its an issue his company has not yet solved.
Key Facts
Evan Hubinger, the Alignment Science Lead at Anthropic, wrote on X that he and his colleagues do “earnestly believe AI could kill all humans,” and he pegged his own estimate at more than 10% in the next decade.Hubinger said the company was “trying its best,” but it does not yet have a plan to solve the issue of “alignment for superintelligence” and is not “clearly on track” to do so.
Hubinger’s post responded to an X thread by another Anthropic researcher, Jacob Coxon, who announced he is resigning from the company over AI safety concerns.
Coxon, who said he has worked on pretraining research at OpenAI and Anthropic, warned that the rival companies were not acting responsibly by “racing straight to self-improving superintelligence and gambling with our lives.”
The departing researcher noted that people building AI believe it could “kill us all by the end of the decade,” and this was not a “marketing stunt.”
- https://www.forbes.com/sites/siladi...ai-could-kill-all-humans-as-researcher-quits/
Anthropic Researcher Quits Over ‘Out-of-Control’ AI Fears
Concerns are rising inside AI labs that competition is pushing tech companies to race toward self-improving models that risk spiraling out of human control
An Anthropic researcher is quitting the artificial-intelligence industry over fears that the lab and its competitors are racing to build systems they won’t be able to control, a sign of mounting safety concerns within top AI companies.
Jacob Coxon, a researcher who specializes in training new AI models by having them consume vast amounts of data, said Tuesday that he is leaving the company because he doesn’t want to participate in an industrywide rush to build AI systems that can improve themselves, worried such systems could spiral out of control and destroy humanity.
- https://www.wsj.com/tech/ai/anthropic-researcher-quits-over-out-of-control-ai-fears-707b7628
The world’s first advanced general‑purpose humanoid robot has completed automated production and autonomously walks off the line.
Last edited:

