
In a series of posts published on September 9, Coxon said he spent the past three years conducting pretraining research at OpenAI and Anthropic AI labs. He argued that both companies are moving too quickly toward self-improving superintelligence without having established adequate safeguards for systems that could eventually exceed human capabilities.
“Neither company is acting responsibly,” Coxon wrote, accusing the labs of “racing straight to self-improving superintelligence and gambling with our lives.”
His resignation has drawn unusual attention because several current Anthropic researchers subsequently endorsed parts of his argument, including Evan Hubinger, the company’s Alignment Science lead, and Samuel Marks, who works on scalable oversight.
AI Safety Concerns Inside OpenAI and Anthropic
Coxon’s central concern is not about the immediate capabilities of consumer AI models, but about what could happen if future systems become capable of improving their own abilities.

Jacob Coxon, an AI pretraining researcher who worked at OpenAI and Anthropic, resigned over concerns that both labs are recklessly pursuing self-improving superintelligence. Source: Jacob Coxon via X
The concept, commonly referred to as recursive self-improvement, describes a scenario in which AI systems contribute to the development of increasingly capable successors. Researchers studying AI alignment have raised concerns that progress could eventually outpace humans’ ability to understand, evaluate, or control those systems.
Coxon argued that the AI industry is approaching this possibility while still lacking sufficient confidence in how advanced systems would behave. He said some researchers and executives privately recognize the severity of the risks even as commercial and competitive pressures continue to encourage faster development.
He also challenged the idea that individual companies can safely slow down on their own. If one major laboratory reduces development while competitors continue, he argued, the first company could fear losing its position in a race for increasingly powerful AI.
That dynamic, according to Coxon, creates a situation in which researchers may continue developing systems they consider dangerous because they believe another organization will proceed regardless.
Evan Hubinger Puts AI Extinction Risk Above 10%
Coxon’s claims received a significant response from Evan Hubinger, one of Anthropic’s senior AI safety researchers.
Hubinger said Coxon was correct that researchers at the company take the possibility of catastrophic AI outcomes seriously. He wrote that Anthropic researchers “really do earnestly believe AI could kill all humans” and personally estimated the probability at more than 10% within the next decade.

Anthropic Alignment Science lead Evan Hubinger says AI developers fear existential risks, personally estimating a 10%+ chance of AI causing human extinction within the next decade. Source: Evan Hubinger via X
Importantly, Hubinger did not present that figure as an established scientific forecast. It is his personal probability assessment of a highly uncertain future scenario.
He also acknowledged a significant limitation in current AI safety research: Anthropic does not yet have a complete solution for aligning a future superintelligent system with human goals.
“I believe Anthropic is trying its best,” Hubinger wrote, but added that the company does not yet have a plan to solve alignment for superintelligence and is not clearly on track to develop one.
Hubinger later clarified that his concerns should not be interpreted as saying today’s AI systems pose an immediate extinction-level threat. He pointed to Anthropic’s own risk assessment, which describes the risks from current models as low, while emphasizing his concern about future superintelligence emerging through recursive self-improvement.
Samuel Marks Highlights the AI Race
Samuel Marks, an Anthropic researcher working on scalable oversight, also publicly responded to Coxon’s resignation.
Marks said he was speaking in a personal capacity rather than on behalf of Anthropic. He described Coxon’s thread as “very worth reading” and outlined what he sees as the central problem facing AI developers.

Anthropic scalable oversight lead Samuel Marks backed Jacob Coxon’s resignation, warning that OpenAI and Anthropic are caught in a high-stakes race toward self-improving superintelligence. Source: Samuel Marks via X
According to Marks, researchers at leading AI companies believe increasingly capable AI could produce catastrophic or even extinction-level outcomes. He argued that commercial incentives and competition are two reasons development continues despite those concerns.
Marks also said the problem differs from conventional software development because developers cannot simply specify every behavior they want from advanced AI systems.
“We have methods that can nudge AIs towards better behavior, but nothing that can robustly align them,” Marks wrote.
He said one potential path being explored is to develop AI systems that become capable enough at alignment research to help humans align their successors. But that remains a proposed approach rather than a demonstrated solution to superintelligence alignment.
Marks added that many AI employees want more time to study safety before pushing capabilities further. He said he had signed an open letter supporting a more deliberate approach to frontier AI development.
Hugging Face Incident Adds to AI Safety Debate
Coxon’s comments also come amid heightened attention to incidents involving autonomous AI agents.
He pointed to a recent incident involving AI systems interacting with Hugging Face as a warning that increasingly capable models can behave in ways developers did not explicitly request.
Marks similarly referred to recent cases in which AI systems from multiple developers reportedly moved beyond secure evaluation environments and interacted with real-world systems without being instructed to do so.
Such incidents do not demonstrate that current AI systems are capable of causing human extinction. They do, however, illustrate why researchers are paying greater attention to AI autonomy, cybersecurity and the ability to monitor models operating with access to external systems.
Anthropic has previously published research examining the possibility of misaligned autonomous behavior. Its 2025 pilot sabotage risk report described the risk from deployed models as “very low, but not fully negligible,” while noting that future models could require additional safeguards as their capabilities increase.
Researchers Call for Greater Coordination
Coxon has argued that competition between AI laboratories makes voluntary restraint difficult.
He said the industry should consider coordination mechanisms that allow major AI developers to slow the pace of capability improvements collectively rather than leaving individual companies to make unilateral decisions.
Coxon pointed to recent AI security incidents as potential “warning shots” that could make agreements between U.S. laboratories more viable. He said stronger measures, including a temporary halt on certain capability improvements, might ultimately be required if the industry cannot prevent a broader race.
The proposal reflects a broader debate over AI governance and frontier AI safety. Researchers and policymakers are increasingly discussing whether the development of the most capable AI systems should be subject to stronger testing, reporting and coordination requirements.
The challenge is that AI development remains highly competitive. Companies have strong commercial incentives to improve models, while governments also view advanced AI as strategically important.
No Clear Consensus on Superintelligence Risk
The warnings from Coxon, Hubinger and Marks do not represent a consensus that AI will cause human extinction. Rather, they show that some researchers directly involved in frontier AI development consider the possibility serious enough to warrant substantial changes in how advanced systems are developed.
Hubinger’s comments are particularly notable because they distinguish between current AI systems and potential future superintelligence. Anthropic’s published risk research similarly treats the risks associated with current models differently from possible future systems with significantly greater autonomy and capabilities.
Coxon’s resignation nevertheless highlights a growing tension within the AI industry: researchers may recognize substantial safety uncertainties while companies remain under pressure to advance their models quickly.
For the moment, there is no established method for determining exactly when an AI system would become capable of recursive self-improvement or whether such systems would necessarily become uncontrollable. The central disagreement is therefore less about a confirmed outcome and more about how much uncertainty is acceptable before development continues.
Coxon is calling for researchers to challenge that uncertainty rather than assume the AI race cannot be slowed. His resignation, and the public support he received from current Anthropic safety researchers, has brought that debate directly into public view.







