Two Anthropic Researchers Leave the Company Over Concerns About the Race for Superintelligence

13.09.2026 9 minutes Author: Newsman

Two researchers who worked on the development and safety of advanced artificial intelligence systems have left Anthropic and publicly warned about the risks of the current AI race. In their view, model capabilities are advancing so quickly that safety and control mechanisms may simply fail to keep pace.

Jacob Coxon and Joe Benton.

One of the most high-profile departures was that of Jacob Coxon. For roughly three years, he worked on research in model pretraining, first at OpenAI and later at Anthropic. In September, Coxon announced that he was leaving the company, explaining that his decision was not about joining a competitor or launching his own startup, but about fundamental concerns over the direction in which the modern AI industry is heading. In his post on X, he said directly that he no longer wanted to take part in the race to build increasingly powerful systems whose consequences, in his view, could be dangerous for humanity.

One of Coxon’s subsequent posts quickly went viral on the platform. In it, the researcher wrote that people directly involved in building today’s AI systems seriously consider a scenario in which artificial intelligence could pose a lethal threat to humanity by the end of the decade. Coxon stressed that he does not see such warnings as marketing hype: according to him, some executives and senior researchers express their concerns far more cautiously in public than they do in private conversations. At the time of the screenshot, the post had received 26.1 million views, 129,000 likes, 15,000 reposts, 10,000 bookmarks, and around 1,400 replies.

In an interview with WIRED, Coxon explained his concerns in greater detail. According to him, leading AI labs are moving toward systems that will not only be able to perform individual tasks, but also help create even more powerful models by automating part of the researchers’ own work. He considers this scenario especially dangerous: as long as new models are built by humans, the pace of progress is at least partly limited by the speed of engineers, scientists, and available computing resources. But if sufficiently powerful AI can independently write code, test architectures, run experiments, and help train the next generations of models, the pace of progress could accelerate sharply.

Jacob Coxon is a former researcher at OpenAI and Anthropic who worked on training advanced AI models.

Coxon Believes Companies Are Taking Too Much Risk

In TechCrunch, Coxon’s position is described in even stronger terms. He believes the problem lies not only in the rapid pace of technological development, but also in the very logic of competition among the world’s largest AI companies. OpenAI, Anthropic, Google DeepMind, and other labs are investing enormous resources in building increasingly powerful models, but each of them risks losing ground if it decides to slow down while competitors continue moving forward. In Coxon’s view, this is exactly what creates a dangerous race in which even people who are seriously concerned about the potential consequences are reluctant to be the first to stop.

Coxon describes the end point of this race as self-improving superintelligence — a hypothetical AI whose intellectual capabilities would significantly surpass those of humans and which could help create even more powerful systems. He described the current situation in especially stark terms, saying that companies are effectively “gambling with our lives.” At the same time, this does not mean Coxon believes today’s ChatGPT or Claude systems are already uncontrollable. His concerns are focused primarily on what could emerge in the future if AI capabilities continue to grow faster than people’s ability to test, constrain, and understand the behavior of such systems.

The Researcher Left Shortly Before Receiving Part of His Anthropic Equity

Coxon’s decision also drew attention because of its financial implications. As Axios reported, he left Anthropic roughly two months before he was due to gain rights to part of his equity in the company. Given Anthropic’s valuation, that stake could potentially have been worth a significant amount, and Coxon himself pointed out that after leaving he no longer had a direct financial interest in the company’s valuation continuing to rise.

This detail made his statement even more notable, because his departure is difficult to explain simply as a desire to change employers or secure a more lucrative position. According to Coxon, the main reason was his unwillingness to continue working on a technology whose risks he considers too high. In other words, his departure was not merely a career move, but a public signal that he disagrees with the current pace and direction of the industry.

Joe Benton Raised Similar Concerns

Coxon was not the only Anthropic employee to leave the company amid similar concerns. Joe Benton, who worked on Anthropic’s safety team and focused on oversight of advanced AI systems, also publicly explained his departure. It is important to understand the timeline correctly: Benton had actually left Anthropic about two weeks before publishing his explanation. In his own post, he said that the attention surrounding Coxon’s statement prompted him to explain his reasons earlier than he had originally planned.

The two researchers’ views overlap in many respects. Benton also believes that leading AI companies are gradually moving toward systems that could significantly accelerate further AI research. If such a system becomes capable of effectively helping to create its own next generations, the pace of technological development could shift from merely very fast to a level that becomes increasingly difficult for humans to keep up with. Like Coxon, Benton is not talking about a science-fiction scenario in which AI suddenly becomes “evil.” His concern is more about a situation in which an extremely powerful system is given a goal but pursues it in ways that developers did not anticipate or can no longer fully control.

Joe Benton is a former researcher on Anthropic’s safety team who worked on AI control and risk assessment.

Benton Did Not Leave the Field of AI Safety

Unlike Coxon, Joe Benton did not decide to step away from artificial intelligence altogether. After leaving Anthropic, he joined METR, an independent organization that evaluates the capabilities and potential risks of advanced AI models. His decision helps explain one of the main criticisms of the current safety system: the companies building the most powerful models are also largely responsible for deciding how safe those models are. Benton believes that this is not enough and that the industry needs stronger independent oversight.

In his view, outside organizations should be able to test the most advanced systems, assess their autonomy, their ability to carry out complex tasks, and any potentially dangerous capabilities. This is especially important in a competitive environment: if Anthropic were to slow down significantly while OpenAI, Google DeepMind, or other companies continued moving forward, it could lose its technological advantage. As a result, all of the major players may acknowledge that the risks are real while still continuing to accelerate, because no one wants to be the first to stop.

Those Who Remain at Anthropic Are Also Talking About the Risks

It is particularly notable that these concerns are not being raised only by former employees. People who continue to work on model safety inside Anthropic have also spoken publicly about the possibility of serious risks. One of them is Evan Hubinger, who works on alignment research — in other words, on making AI systems behave in ways that remain consistent with human goals and intentions.

Hubinger has not left Anthropic, but he has also acknowledged the possibility of very serious consequences if future systems develop in an uncontrolled way. The fact that current researchers are expressing similar concerns makes the debate much broader than the story of two employees who chose to resign. At the same time, there is no full scientific consensus on how realistic a rapid emergence of superintelligence or a loss of control over AI actually is. Some researchers consider such forecasts too speculative, while precise estimates of the probability of a global catastrophe are effectively impossible to verify.

Even Anthropic’s Leadership Says the Industry Needs to Be More Cautious

The situation becomes even more striking when viewed against the position of Anthropic’s own leadership. The company’s CEO, Dario Amodei, has repeatedly warned that the capabilities of advanced models could develop faster than the mechanisms designed to control and regulate them. In The Guardian, his position is described as a call for a more cautious approach to AI development and stronger independent testing of the most powerful models.

Amodei has also argued for broader coordination between companies and governments. The logic is that a single lab cannot solve the problem on its own if other players continue expanding the capabilities of their systems without similar restrictions. This creates a fairly contradictory picture: Anthropic, like its competitors, continues to invest enormous resources in building increasingly powerful AI, while some current and former employees openly warn that the pace of development could itself become a problem.

The Main Risk Is Not That AI Will Suddenly “Come Alive”

Discussions about AI risks can easily drift into science-fiction imagery in which a machine suddenly becomes conscious and decides to destroy humanity. But researchers such as Coxon and Benton are generally not talking about that kind of scenario. Their concern is more about the emergence of systems that become so capable and autonomous that they can carry out complex tasks for long periods of time without constant human supervision.

If such models gain the ability to write code independently, conduct scientific research, carry out cyber operations, interact with financial systems, or control real-world infrastructure, even a small error in how a goal is defined could have consequences far more serious than a wrong answer from a modern chatbot. That is why the central question in this debate is not “will AI become evil?” The more important question is whether humans will be able to build reliable control mechanisms for systems that could eventually become smarter, faster, and more autonomous than their own developers.

The departures of Jacob Coxon and Joe Benton do not mean that a catastrophic scenario is inevitable. But their decisions show that, for some of the people working directly with the most powerful AI systems, the possibility of losing control over future AI has already moved beyond the realm of science fiction.

Subscribe
Notify of
0 Коментарі
Oldest
Newest Most Voted
Found an error?
If you find an error, take a screenshot and send it to the bot.