Anthropic Alignment Lead Issues Warning About AI Killing Humans As Researcher Resigns

Topline

A senior researcher who leads Anthropic’s alignment efforts said Tuesday that many in the company believed AI could wipe out humanity and warned that the company is not on track to solve the issue of aligning AI’s goals with humanity’s, despite its efforts, after another researcher quit the company, accusing it of not acting responsibly.

Key Facts

Evan Hubinger, the Alignment Science Lead at Anthropic, wrote on X that he and his colleagues do “earnestly believe AI could kill all humans,” and he pegged his own estimate at more than 10% in the next decade.

Hubinger said the company was “trying its best,” but it does not yet have a plan to solve the issue of “alignment for superintelligence” and is not “clearly on track” to do so.

Hubinger’s post responded to an X thread by another Anthropic researcher, Jacob Coxon, who announced he is resigning from the company over AI safety concerns.

Coxon, who said he has worked on pretraining research at OpenAI and Anthropic, warned that the rival companies were not acting responsibly by “racing straight to self-improving superintelligence and gambling with our lives.”

The departing researcher noted that people building AI believe it could “kill us all by the end of the decade,” and this was not a “marketing stunt.”

crucial quote

In a post following up his dire warning, Hubinger cited Anthropic’s latest risk report and said the threat posed by present models is low and added: “What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought.”

What Do We Know About The Anthropic Researcher’s Resignation?

The Wall Street Journal first reported Coxon’s departure from Anthropic over concerns about the push to build self-improving AI systems that could threaten humanity. Coxon told the Journal that he believes the world is on track for a “lot of the most aggressive of these scenarios where by the end of next year things could be out of control already.” In his X posts, Coxon compared working at OpenAI and Anthropic, noting that staffers at the ChatGPT-maker have not “deeply internalized the civilizational stakes.” He said the stakes were “well-understood” at Anthropic, but the company was “locked in a race to get there first” as they believe “no one else will act responsibly, so they must do it themselves, despite the risk.”

Key Background

The dire warning and resignation come on the backdrop of a push by some leading AI researchers and executives for a slowdown in advanced AI development. A statement titled “Pacing the Frontier” was signed by several top AI figures in July, including Anthropic co-founders Dario Amodei and Jared Kaplan, OpenAI Chief Scientist Jakub Pachocki, Meta AI chief scientist Shengjia Zhao, and others. The statement said that to “realize AI’s potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight,” but warned that companies face intense competitive pressure to do so unilaterally. The statement urged the U.S government to support an “international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” In a blog post on Sunday, Pachocki echoed these warnings, noting that he believes “no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”

further reading

Anthropic Researcher Quits Over ‘Out-of-Control’ AI Fears (Wall Street Journal)

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top