Anthropic researcher resigns, says AI labs are racing recklessly toward superintelligence
A person posting on X under the handle @hilbertspaess, identifying himself as Jacob Coxon, announced in a seven-tweet thread on September 9, 2026 that he resigned from Anthropic that day. He says he spent the last three years doing pretraining research across both OpenAI and Anthropic combined, and that neither company is acting responsibly: both are, in his words, racing straight to self-improving superintelligence and gambling with everyone's lives.
He argues the technology should not be underestimated. He expects systems that can hack anything, transform entire fields overnight, and accumulate real power and resources, and says progress toward that point is not slowing.
His central claim is about belief, not just capability: the people building this technology, he says, earnestly believe it could kill everyone by the end of the decade, and this is not a marketing stunt. He says many executives and senior researchers soften their public language to sound reasonable, while privately expressing the same fear.
He draws a distinction between the two labs. At OpenAI, he says, many people have not deeply internalized the civilizational stakes. At Anthropic, he says the stakes are well understood, but the company is locked in a race to get there first: its leadership believes no one else will act responsibly, so Anthropic must do it itself despite the risk.
He calls the decision to accept this race and enter what he terms the "endgame" a hubristic gamble that should not be launched from a private company's Slack, and argues that trying to speedrun alignment should demand extraordinary confidence that no better path exists.
He says he remains optimistic about the possibility of coordination between labs. He cites the Hugging Face attack as a warning shot that has made pacing agreements between US labs more viable, though he does not describe what the attack involved. He adds that he does not feel the world is on track to prevent a global AI race, and that preventing it may require costly measures such as a temporary ban on improving model capabilities.
He closes by addressing lab researchers directly, asking them to weigh what the next few years will actually feel like: whether to launch a superintelligent reinforcement-learning run without a rigorous understanding of the resulting system's mind, keep working because development is happening anyway, or use this moment to push for different conditions.
Key facts
- A person posting as @hilbertspaess, identifying himself as Jacob Coxon, says he resigned from Anthropic on September 9, 2026, after three years doing pretraining research across both OpenAI and Anthropic combined.
- He says both companies' senior people privately believe AI could kill everyone by the end of the decade while publicly softening that message for the press.
- He distinguishes the labs: OpenAI, he says, has not deeply internalized the stakes; Anthropic understands them but is racing to get there first because its leadership believes no one else will act responsibly.
- He cites the Hugging Face attack as a warning shot that has made pacing agreements between US labs more viable, and floats a temporary ban on improving model capabilities as a possible costly remedy.
- He closes by urging other lab researchers to decide between continuing as before or pushing for different conditions.
Why it matters
Public resignations citing existential-risk concerns from people claiming direct pretraining experience at both OpenAI and Anthropic are rare. Whether or not every detail can be checked, the thread surfaces a claim that matters on its own terms: that senior people inside frontier labs privately voice fears about AI risk they do not voice publicly, and that Anthropic's own account of its strategy, as related here, is to race toward superintelligence because it doubts anyone else would do so responsibly.
Who it affects
The claims concern OpenAI and Anthropic specifically, and by extension anyone tracking the internal culture and incentives at frontier AI labs: employees weighing whether to stay, policymakers and researchers interested in inter-lab coordination, and the wider public debate over how seriously AI safety risk is treated inside the companies building the technology.
How to use it
This is a personal essay, not a product or a paper, so there is nothing to buy or install. Its practical value is as a signal to watch two things going forward: whether pacing agreements between US labs actually materialize, as the author says the Hugging Face attack has made more likely, and whether other researchers make similar public departures for similar reasons.
How solid is it
The account rests entirely on one person's word, posted from an X handle, with no independent confirmation offered in the thread itself that its author worked at either company. Replies beneath the thread reportedly question the claim, noting the account is recent, though that is discussion rather than part of the source. No names, teams, managers, or details of the cited Hugging Face attack are given, which limits what can be checked.
Risks and caveats
Treat this as one person's self-reported account and stated opinions, not a verified report: the employment history, the internal characterizations of Anthropic's and OpenAI's leadership, and the description of the Hugging Face attack are all asserted, not documented. The thread names no other individuals and does not claim any other named employee has resigned for the same reasons.
“They are racing straight to self-improving superintelligence and gambling with our lives.”
— Jacob Coxon (@hilbertspaess)