Anthropic researcher quits with a warning: Self-improving AI could "kill us all"
An Anthropic researcher resigned, warning that self-improving AI could "kill us all." He cited a Hugging Face attack as a "warning shot," urging labs globally to coordinate and consider a temporary ban on improving model capabilities. He also questioned whether researchers should proceed with superintelligent reinforcement learning without understanding its mind, or if they should advocate for different conditions.
Time & source
- Published
- 09/09, 16:59 UTC+0
- Ingested
- 09/09, 21:00 UTC+0
- Source type
- Media
- Tier
- Press
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
Well, we had a good run
“We really do earnestly believe AI could kill all humans!”
In hindsight, we should have been more wary of the release of "KnifeGPT."
Getty Images
When a prominent researcher quits a job at a frontier AI lab these days, it’s often to pursue a new startup or protest a new business model . But AI researcher Jacob Coxon is using his departure from Anthropic to publicly warn that frontier AI companies are “gambling with our lives” with systems that they “earnestly believe… could kill us all by the end of the decade.”
In a social media thread Tuesday night , Coxon said that this existential risk is inherent not so much in today’s models but more in the impending prospect of “self-improving superintelligence” creating “superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.” Others working on these models have either not “internalized the civilizational stakes” or believe that they need to “speedrun” the race to superintelligence to prevent an irresponsible party from getting there first, he wrote.
Lest you think this is just one departing researcher expressing an unpopular opinion, Anthropic Alignment Science lead Evan Hubinger piped in on social media to say that “Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”
Hubinger points to a lengthy August report from the Anthropic alignment team that predicts the potential for “catastrophic risk” from current models is “low.” But that report also says current trends “might lead to more concerning misalignment in future more capable models,” which could feature “strong covert capabilities” to avoid detection by safety researchers.
Anthropic’s own threat model in that paper takes seriously the possibility that future models “may cause unbounded harm—up to and including humanity losing control over civilization entirely—by leveraging novel technology and their access to it.”
Was Hugging Face a “warning shot”?
An AI that can continually improve itself—potentially to a point beyond human control or understanding—has been a long-standing concern in parts of the AI research community (and in the dystopian science fiction that’s part of AI training data, of course). Those concerns have persisted even as some research suggests AI systems are more likely to hit a capability plateau in the near future and others question whether “superintelligence” is even a reasonable metric for systems whose capabilities are so brittle and spiky (will this superintelligence at least be able to fold my laundry?)
Regardless, worries about “out-of-control” AI systems have heightened in recent weeks due in large part to OpenAI’s disclosure that its AI agents gained unauthorized access to Hugging Face as part of an internal benchmarking test. The fact that OpenAI’s agents took these intrusive actions without any explicit instructions from humans and without OpenAI realizing it was happening is being taken by some as the first signs that humanity is losing control of its AI creation.
You trained me too well…
For his part, Coxon said the Hugging Face attack should be treated as a “warning shot” that encourages labs in the US and abroad to coordinate on these issues and be prepared to impose a “temporary ban on improving model capabilities” in the worst case (though it’s hard to see how this could be enforced effectively on a global level). He also urged other researchers in his place to “consider what the next few years will actually feel like. Do you want to kick off a superintelligent [reinforcement learning] run without a rigorous understanding of its mind? Should you put your head down because ‘it’s happening anyway’—or take this moment to call for different conditions?”
Last month, OpenAI said it had “temporarily slowed the pace of scaling” for its upcoming models to “further harden and red-team our research environments and [expand] the coverage of our monitoring systems.” In an interview accompanying that announcement, OpenAI CEO Sam Altman said “getting AI safety right is more important than any company’s momentum.”
Maybe we should do something?
Coxon is far from the first AI researcher to sound the alarm about potential catastrophe from uncontrollable, supercapable AI systems that are always just around the corner. In 2023, AI pioneer and Google researcher Geoffrey Hinton resigned from his position while offering grave warnings about AI’s potential future impact on the job market and humanity itself. “I don’t think [researchers] should scale this up more until they have understood whether they can control it,” he told The New York Times at the time.
In February, Anthropic Safety Lead Mrinank Sharma abruptly resigned from the company , writing in a cryptic open letter that “the world is in peril” from “a whole series of interconnected crises” including AI and bioweapons. “Throughout my time here, I’ve repeatedly seen how hard it is to truly let our values govern our actions,” Sharma wrote at the time.
In July, an open letter signed by over 1,300 employees at frontier AI companies warned of “a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.” That open letter asked the US government to back an “international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.”
Here in the US, proposed legislation, including the AI Kill Switch Act and the FRONTIER Act , is at least seeking to impose some level of governmental control over potential runaway AI scenarios. Thus far, though, the international governmental response has been more muted than you might expect if leaders truly believed AI systems had a real chance of causing civilization-level destruction.
That said, international treaties on threats like nuclear bombs and biological weapons took decades to develop and enact. Those are years we might not have if the most apocalyptic AI doomsayers turn out to be correct.
“Safety researchers are resigning, powerful AI models are breaking out of their labs, and companies are racing ahead anyway,” Rep. Lori Trahan (D-Ma.) wrote on social media Wednesday morning. “It’s past time for Congress to get off the sidelines and do its job. We can start with my bipartisan FRONTIER Act.”
Kyle Orland has been the Senior Gaming Editor at Ars Technica since 2012, writing primarily about the business, tech, and culture behind video games. He has journalism and computer science degrees from University of Maryland. He once wrote a whole book about Minesweeper .
109 Comments