Dilshod.dev

Why Anthropic Researchers Are Walking Away

A 27-year-old pretraining researcher quit Anthropic this week, saying the industry is 'gambling with our lives.' He is the second senior departure this year. What the pattern actually tells us — and what it doesn't.

by Dilshod Abdullayev7 min read

On September 8, Jacob Coxon told the Wall Street Journal he was leaving Anthropic — and leaving the field entirely. He is 27. He spent the last three years on pretraining research, first at OpenAI, then at Anthropic. His resignation post on X was brief and unambiguous: neither company is acting responsibly, both are racing toward self-improving superintelligence, and both are gambling with our lives.

It would be easy to file this under "AI doomer quits, film at eleven." I think that reading misses what makes it interesting. The interesting part is not the warning. It's who is giving it, and what he was building when he stopped.

Pretraining is not the safety team

Most public AI criticism comes from outside the building — from academics, journalists, policy people, or from safety teams whose entire job is to worry out loud. Coxon was none of those. He worked on pretraining.

Pretraining is the phase where a model is trained from scratch on an enormous corpus of text. It is the most expensive, most compute-hungry, most foundational part of building a frontier model. Everything downstream — fine-tuning, alignment, the product you actually talk to — is shaping something that pretraining already made. If you work on pretraining, you are not adjacent to the core of the company. You are the core.

That's why this particular resignation lands differently than a broadside from a critic. A critic is reasoning about a system from the outside. Coxon was reasoning about it from inside the loop, with access to the internal numbers, and he concluded that he did not want his name on the next iteration.

The claim he's actually making

His specific fear is worth stating precisely, because the headlines flatten it.

Coxon told the WSJ that we are tracking toward the more aggressive scenarios, and that by the end of next year things could already be out of control. What he means by "out of control" is a system capable of compromising virtually any computer system, transforming entire fields of human activity in a very short window, and accumulating real-world resources and influence.

The load-bearing phrase in all of this is self-improving. Today, humans design each generation of model. The concern is what happens when a model becomes genuinely good at the research that produces the next model. At that point the loop closes: each generation shortens the time to the following one, and the humans who were supposed to be steering become the slowest component in their own process. You don't need any science-fiction assumption about consciousness or malice for this to be worrying. You only need the loop to be faster than the institutions reviewing it.

Here I want to be careful, because this is where reporting usually stops being reporting. The timeline is a prediction, not a fact. "By the end of next year" is one researcher's forecast. Forecasts from smart, well-positioned people have been wrong in both directions throughout the history of this field. What is verified: he worked on pretraining at both labs, he resigned, and he said these words publicly. What is not verified: that he is right.

The second data point

This is where it stops being one person's story.

In February of this year, Mrinank Sharma resigned as head of Anthropic's Safeguards research team. His resignation letter said "the world is in peril" — not from AI alone, he wrote, but from a set of interlocking crises happening at once. He also said something more specific and, for an engineer, more recognizable: that the safety team constantly faced pressure to set aside what mattered most. The letter reached close to a million views on X.

So within seven months, Anthropic lost the head of a safety research team and a core pretraining researcher. Different roles, different arguments, one shared structure: neither of them said the company was malicious. Both said the pace of the race outruns the pace of responsibility.

That distinction matters. "This company is evil" is a claim about intent, and it's usually wrong. "This organization's incentives are moving faster than its ability to check itself" is a claim about structure — and structural claims are the ones worth taking seriously, because they don't require anyone to be a villain.

Why the timing is not an accident

Anthropic is in a pre-IPO period. That single fact reframes everything above.

Going public means outside investors, quarterly expectations, and permanent pressure to ship faster and monetize harder. Anthropic's entire brand — its founding story, its public positioning, its CEO's repeated calls for regulation to force the industry to slow down — is built on being the lab that puts safety first. Two insiders publicly disputing exactly that story, months before an IPO, is not a technical problem. It's a credibility problem, and credibility is most of what a safety-first brand actually owns.

There is a version of this where the departures are the system working: people with real objections leaving loudly, in public, in a country where they're free to do so. That's genuinely better than silent attrition. But it also means the objection was not resolvable from the inside — and that's the part worth sitting with.

What this means if you build software

You are probably not training frontier models. I'm not either. Here's what I think transfers anyway.

The pattern Sharma described — the team responsible for caution facing constant pressure to set aside what matters most — is not exotic. It is the most ordinary failure mode in software. It is the security review that becomes a checkbox before a launch date. It is the migration plan that gets compressed because the quarter is ending. It is the test suite that gets marked flaky instead of fixed. The scale differs by many orders of magnitude; the mechanism is identical. Speed is legible to management and caution is not, so caution is what gets traded away first, quietly, without anyone deciding to.

The practical version for the rest of us: notice when your own "we'll handle it after launch" list stops getting shorter. That list is a measurement. If it only ever grows, your organization has already made a decision about the tradeoff between speed and care — it just hasn't said so out loud.

What I'd hold lightly

Two things.

First, resignation warnings are a genre now, and genres attract performance. That's not an accusation against Coxon or Sharma specifically — I have no basis for one — but a warning delivered on X, amplified by a WSJ interview, is a message with an audience, and the audience shapes the message. Read it as testimony, not as a scientific result.

Second, insiders are not automatically right. Being close to a system tells you how it works; it does not make you a good forecaster of where it goes. Some of the least accurate predictions in tech have come from people with the best possible seats.

What I'd hold firmly is narrower: two people with direct access to Anthropic's core work left within seven months, both citing pace over safety, one of them from the team whose job was safety itself. That is not proof of anything. It is a signal, and signals from inside are worth more than opinions from outside.

Sources: Wall Street Journal (via Free Press Journal), TradingKey, Forbes, Semafor.

Related reading

aicareer

The Missing First Rung

Stanford tracks the payroll of five million Americans a month. Employment for 22-to-25-year-olds in AI-exposed jobs sits nineteen percent below where it should be. Nobody is being laid off — they are simply not being hired. Those are entirely different problems.

8 min read
aiengineering-culture

The AI Wrote It. You Still Own It.

Generating code got cheap. Understanding it did not. Review is now the bottleneck skill, and reviewing generated code is harder than reviewing a colleague's.

11 min read
aiagents

Context Engineering Is the New Bottleneck

Prompt engineering was about wording. Context engineering is about deciding which tokens deserve a place in a limited window — and it is now the skill that separates agents that work from agents that almost work.

6 min read