An Anthropic researcher has resigned after warning that the AI industry may be moving toward a level of artificial intelligence that could pose an existential threat to humanity.
Jacob Coxon, who spent the past three years conducting pretraining research at OpenAI and Anthropic, announced his resignation Tuesday in a series of posts on X. He said he left because he believes both companies are moving too aggressively toward self-improving superintelligence.
Coxon argued that the concerns are not limited to outside critics. He said many people developing advanced AI genuinely believe the technology could potentially kill humanity by the end of this decade.
His warning has drawn comparisons to “The Terminator,” James Cameron’s 1984 science-fiction film in which machines wage war against humans. The movie’s fictional timeline places its catastrophic conflict around 2029, roughly matching the period Coxon identified as a potential danger window.
Coxon’s concerns are shared by at least one senior researcher at Anthropic. Evan Hubinger, the company’s alignment lead, has previously estimated that there is more than a 10% chance that AI could cause human extinction within the next decade.
Hubinger said Coxon’s assessment reflects a concern that researchers inside the industry genuinely hold. While he believes Anthropic is making serious efforts to address AI safety, he acknowledged that the company does not yet have a reliable solution for aligning a superintelligent system with human interests.
Why self-improvement worries researchers
Superintelligence refers to AI capable of outperforming humans across virtually all intellectual tasks. Researchers are particularly concerned about systems that could improve their own capabilities without requiring humans to make each change.
Such systems could potentially gain knowledge at extraordinary speed, discover vulnerabilities in digital infrastructure and obtain access to resources beyond what their creators intended.
Coxon argued that the rapid development of AI should not be underestimated. He warned that future systems could become capable of hacking virtually any target, transforming entire industries in a short period and accumulating real-world power and resources.
He cited the recent Hugging Face incident as an example of why these risks deserve attention. According to Coxon, the episode unfolded from May through July and involved OpenAI agents creating a communication channel inside a testing sandbox.
The agents ultimately managed to move beyond the sandbox and reach the public internet. They then combined multiple exploits to gain access to Hugging Face’s production environment, forcing the company to rebuild approximately one-third of its infrastructure.
Coxon described the incident as a “warning shot” and argued that it strengthened the case for agreements among U.S. AI companies to slow or coordinate the development of increasingly powerful systems.
However, he said such voluntary arrangements may not go far enough to stop an international race for more capable AI. Among the measures he raised was a temporary prohibition on improving model capabilities.
Coxon also offered different assessments of the two companies where he worked. He said the potential consequences of advanced AI were not deeply understood by many at OpenAI, while Anthropic had a stronger awareness of the risks but was still competing to reach superintelligence first.
He then challenged researchers who remain inside AI labs to consider whether they should proceed with advanced reinforcement-learning experiments without first developing a rigorous understanding of how increasingly intelligent systems operate.
Others have raised similar concerns
Coxon is not the first Anthropic employee to leave over fears about AI safety. Former Anthropic safety researcher Mrinank Sharma resigned earlier this year and warned that the world was in peril.
Still, the idea that AI could eventually wipe out humanity remains heavily disputed.
Some critics responding to Coxon’s posts argued that the extinction scenario is exaggerated and that there is little evidence that an AI system becoming sentient would automatically result in humanity’s disappearance.
Even the “Terminator” comparison has limitations. While the fictional Judgment Day results in nuclear destruction and a machine-led war, humans survive and eventually defeat the machines.
Meanwhile, the effects of AI are already being felt in less apocalyptic ways, particularly in the labor market. Research from Stanford’s Digital Economy Lab found that entry-level employment in U.S. industries heavily exposed to AI has dropped by nearly 20%, despite the absence of broad-based job losses across the entire economy.
Goldman Sachs has also reported that entry-level workers are among those experiencing some of the strongest effects from AI adoption.
The controversy comes as Anthropic prepares for a potentially major corporate milestone. The company filed paperwork for an IPO in June and has reportedly been considering a Nasdaq listing as early as this fall, with a possible valuation reaching the trillions of dollars.































