An Anthropic Researcher Just Quit, Warning the AI Race Is Now the Real Danger
Jacob Coxon spent three years training frontier models at OpenAI and Anthropic. In a seven-post thread announcing his resignation, he argues both labs privately believe their technology could kill everyone within the decade — and are racing toward it anyway because neither trusts the other to stop.
Jacob Coxon has spent the last three years inside the two labs building the world's most capable AI systems — first at OpenAI, then at Anthropic, working on pretraining, the process that produces a model's raw capabilities before any safety fine-tuning is layered on top. On September 9, 2026, he resigned from Anthropic and explained why in a seven-post thread on X, one of the most direct public accounts yet of what people inside frontier labs say to each other versus what they say in public.
His opening line set the tone: "Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives."
"The same people express fear privately"
The most striking claim in Coxon's thread isn't about capability timelines — it's about what he says researchers and executives actually believe versus what they say for the record:
"The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible — but I hear the same people express fear privately. No other human activity poses this level of danger."
This is a pointed claim: that public statements from AI labs are, if anything, understated relative to internal belief — the opposite of the usual accusation that labs hype risk for attention or regulatory capture.
Why build something you fear?
Coxon anticipates the obvious rebuttal — if researchers truly believe this, why keep building it — and answers it differently for each company he worked at:
"A common response is 'if they truly believe this, why are they still building it?' At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first — they believe no one else will act responsibly, so they must do it themselves, despite the risk."
In other words: not denial, but a prisoner's dilemma. Anthropic's own founding rationale has long been that a safety-focused lab needs to be at the frontier to influence how the technology develops — Coxon is describing that logic curdling into the very race it was meant to prevent. He calls the premise underneath it a "hubristic gamble":
"Accepting this race and entering the 'endgame' is a hubristic gamble that should not be launched from a private company's Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available."
What's actually at stake, in his view
Coxon is not vague about capability. He argues the systems now being trained will not stay comparable to today's tools for long:
"Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing."
That combination — offensive cyber capability, real-world resource acquisition, and a self-improvement loop — is the specific scenario alignment researchers describe as losing a meaningful "off switch": today it's trivial to shut down a datacenter; it is not obvious that remains true once such systems are embedded in infrastructure, robotics, or corporate decision-making.
A narrow window for coordination
Coxon isn't purely fatalistic. He points to the aftermath of the Hugging Face breach as evidence that pacing agreements between U.S. labs are more achievable than they looked a year ago — but he's explicit that the current trajectory isn't good enough:
"I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don't feel like we're on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities."
His closing message was aimed squarely at the people still inside those labs:
"If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because 'it's happening anyway' — or take this moment to call for different conditions?"
Not an isolated data point
Coxon's departure lands alongside separate, corroborating signals from inside Anthropic. Evan Hubinger, who leads the company's alignment science team, has previously put the odds of an AI-driven catastrophe at ">10% within the next decade" while acknowledging the field "does not yet have a plan to solve alignment." Coxon's exit is among the first public resignations explicitly framed around AI safety fears — an earlier Anthropic safety researcher left the company this year to study poetry instead.
As of publication, neither Anthropic nor OpenAI has issued a public response to Coxon's thread, which had drawn more than 70 million views within hours.
Sources: Jacob Coxon (@hilbertspaess on X); The Wall Street Journal