Anthropic Researcher Jacob Coxon Resigns With Warning That AI Race Is ‘Gambling With Our Lives’

Anthropic

An Anthropic AI researcher has resigned from the company with a stark warning about where the artificial intelligence industry may be heading, accusing leading AI labs of racing toward self-improving superintelligence without adequate safeguards.

Jacob Coxon, who says he spent the past three years conducting pretraining research at OpenAI and Anthropic, announced his resignation from Anthropic on X on September 9.

“Neither company is acting responsibly,” Coxon wrote. “They are racing straight to self-improving superintelligence and gambling with our lives.”

His comments have reignited a long-running debate inside the AI community: whether companies developing increasingly powerful models can move quickly enough to remain competitive while also ensuring that future systems remain controllable.

Coxon’s claims are his own and should not be treated as evidence that Anthropic or OpenAI has privately reached an agreed conclusion that superintelligent AI will cause human extinction. However, his position is notable because it comes from someone who says he has worked directly on pretraining at two of the world’s most prominent AI companies.

What did Jacob Coxon say about Anthropic and OpenAI?

Coxon’s resignation statement goes well beyond criticism of corporate culture.

He argues that both companies are moving toward systems capable of improving their own capabilities, potentially creating a feedback loop in which increasingly capable AI helps develop the next generation of even more capable AI.

According to Coxon, the companies are not simply pursuing better chatbots or productivity tools. He believes they are moving toward a fundamentally different stage of AI development.

He wrote that future systems could become “superhuman,” potentially capable of hacking computer systems, accelerating scientific research and accumulating access to resources and real-world power.

The most dramatic part of his warning concerned the views he says exist within the industry.

“The people building AI earnestly believe that it could kill us all by the end of the decade,” Coxon wrote.

That is a serious allegation, but it is important to distinguish between Coxon reporting his perception of private discussions and independently verified evidence that AI executives collectively hold that belief.

What is self-improving superintelligence?

The idea at the center of Coxon’s warning is often described as recursive self-improvement.

In simple terms, imagine an AI system that is capable of conducting research into AI itself. If that system can identify ways to make its own algorithms, training methods or reasoning abilities better, developers could use those discoveries to create a stronger version.

A stronger system could then be better at AI research, potentially accelerating the next improvement.

That creates a hypothetical cycle:

AI system → improves AI research → better AI system → better AI research → increasingly capable AI

The concern among some AI-safety researchers is that this process could eventually move faster than human institutions can understand or regulate.

It is important to stress that this remains a future scenario, not an established description of current AI systems. Today’s frontier models can perform sophisticated coding, reasoning and research tasks, but there is no publicly demonstrated system that autonomously and indefinitely improves itself into a superintelligence.

Why is recursive self-improvement such a concern?

The concern is not simply that a very intelligent AI could make mistakes.

It is that a system capable of improving AI development could potentially accelerate its own progress.

Human researchers work within relatively slow institutional cycles. Experiments require computing resources, evaluation, review and deployment decisions. An AI system operating at machine speed could, in theory, perform parts of this process much faster.

That could create several problems.

The alignment problem

An AI system could become extremely capable while still pursuing goals that do not reliably match human intentions.

This is known as the AI alignment problem.

The difficulty becomes more significant as systems become more autonomous. A chatbot answering questions is one thing; an AI agent capable of conducting research, writing and executing code, operating computer systems and making long sequences of decisions is another.

The control problem

Researchers also face a basic question: if a future AI becomes substantially more capable than its operators, how can humans reliably stop or redirect it?

The answer is not obvious.

Traditional software can generally be switched off because its capabilities remain bounded by the programs and infrastructure humans control. A sufficiently autonomous AI system could potentially interact with networks, tools and other systems in ways that make control considerably more complicated.

Again, this is a risk scenario—not evidence that current AI systems have developed such capabilities.

Why would Anthropic continue developing increasingly powerful AI?

Coxon acknowledges the apparent contradiction.

Anthropic has positioned itself as an AI-safety-focused company and has publicly emphasized responsible development. Yet it is also competing directly with OpenAI, Google and other companies to build increasingly capable frontier models.

Coxon argues that Anthropic understands the risks but is nevertheless caught in a competitive race.

“At Anthropic, the stakes are well-understood, but they are locked in a race to get there first,” he wrote.

His characterization is an allegation about the company’s internal decision-making, rather than an independently established fact.

The broader economic incentive, however, is straightforward: companies that build more capable AI systems can potentially gain enormous advantages in software development, research, automation and other industries.

That creates a difficult strategic problem.

If one company slows development because it believes its competitors will continue, it may fear falling behind while doing little to reduce the overall pace of AI progress.

Coxon says AI development should not happen inside private companies

One of Coxon’s strongest arguments concerns governance.

He questioned whether decisions about potentially transformative AI systems should effectively be made by a handful of private companies and their employees.

His concern reflects a larger debate over who should decide how powerful AI systems are developed and deployed.

Possible approaches include:

Coxon suggested that stronger agreements between major AI laboratories could be necessary and raised the possibility of temporarily banning further capability improvements under extreme circumstances.

Such proposals remain controversial.

Critics of strict pauses argue that slowing responsible companies may simply allow less transparent actors to move ahead. Supporters counter that an uncontrolled race can create incentives for companies to prioritize speed over safety.

What does “AI could kill us all” actually mean?

The phrase is deliberately alarming, but it encompasses several different hypothetical risks.

Researchers concerned about advanced AI have discussed scenarios involving:

Misalignment: An advanced system pursues objectives that conflict with human interests.

Cybersecurity: Highly capable AI could potentially make sophisticated cyberattacks easier to execute.

Biological risks: Advanced AI could lower barriers to certain forms of biological research or misuse.

Autonomous systems: AI could increasingly act independently rather than simply provide information to humans.

Loss of control: A sufficiently capable system could become difficult to monitor, constrain or shut down.

These scenarios vary considerably in probability, timeframe and technical feasibility. There is no scientific consensus that any one of them will occur.

That uncertainty is precisely why the debate remains so contentious.

Anthropic has already warned about increasingly capable AI

Coxon’s concerns do not emerge in a vacuum.

Anthropic has itself published research and safety assessments examining risks associated with increasingly capable AI models. The company has also developed internal frameworks intended to identify and manage risks from advanced models.

Researchers associated with Anthropic have separately examined scenarios involving AI systems that could become more autonomous and capable of contributing to AI research.

Evan Hubinger, an Anthropic researcher known for his work on AI alignment, has previously discussed the possibility that future AI systems could behave in ways that are difficult to control and has explored questions surrounding deceptive or strategically misaligned AI.

These discussions do not mean Anthropic believes catastrophe is inevitable. Rather, they demonstrate that the risks Coxon is raising are already part of serious research within the field.

Is AI already capable of improving itself?

Not in the science-fiction sense implied by the phrase “self-improving superintelligence.”

Modern AI systems can assist with coding, generate training data, evaluate outputs and help researchers conduct experiments. AI can therefore participate in parts of the development process.

But there is a major difference between AI-assisted development and a system independently initiating an open-ended cycle of self-improvement.

The latter would require substantially more autonomy, reliable long-horizon planning, strong research capabilities and the ability to implement and evaluate improvements without humans controlling each major stage.

Researchers disagree over how close current systems are to that threshold.

Why Coxon’s resignation matters

Employees leaving technology companies over ethical disagreements is not new. What makes Coxon’s departure notable is the combination of his claimed experience at two major AI labs and the severity of the warning he attached to his resignation.

He is effectively arguing that the industry’s central competitive dynamic is itself becoming a safety problem.

If companies believe that the first organization to achieve a major capability breakthrough will gain enormous economic or strategic advantages, each company has an incentive to continue investing—even if its researchers believe the collective outcome could be dangerous.

That is a classic coordination problem.

No individual company can necessarily solve it alone.

What happens next for the AI industry?

Coxon called on AI researchers to question whether continuing toward increasingly powerful systems simply because “it’s happening anyway” is a sufficient justification.

That question is likely to become more important as AI companies develop systems capable of increasingly autonomous coding, research and decision-making.

The critical issue is not whether AI development should stop forever.

It is whether the safeguards surrounding frontier AI can improve as quickly as the technology itself.

For now, Coxon’s warning remains one researcher’s assessment rather than proof of an imminent AI catastrophe. But it highlights a real tension inside the industry: the same companies investing heavily in AI safety are also competing to build more powerful systems.

That tension is unlikely to disappear.

And if AI eventually becomes capable of substantially accelerating AI research itself, the question of how fast is too fast may become much more consequential than the question of which company gets there first.

TL;DR

Jacob Coxon, who says he spent three years conducting AI pretraining research at OpenAI and Anthropic, has resigned from Anthropic and accused both companies of racing toward self-improving superintelligence without sufficient safeguards.

He warned that some people building advanced AI privately believe the technology could pose an existential threat within the decade. Coxon called for greater coordination among AI companies and suggested that extreme measures, including temporarily halting capability improvements, could eventually be necessary.

His claims have not been independently verified and should not be interpreted as evidence that catastrophe is inevitable. But his resignation adds another insider voice to the growing debate over whether AI development is advancing faster than safety research and governance can keep up.

Exit mobile version