Jacob Coxon Accuses Anthropic of Gambling With Our Lives

Jacob Coxon unclipping his climbing rope while the roped team keeps pushing to the summit

Jacob Coxon, 27, announced overnight that he had resigned from Anthropic after three years of pretraining research at OpenAI and then Anthropic. He accuses both labs of racing toward a self-improving superintelligence while gambling with everyone’s lives. Under ninety minutes later, the person who leads alignment work at Anthropic publicly told him he was right.

Key Takeaways

  • A pretraining researcher leaves Anthropic and publishes his resignation on his own account
  • He targets OpenAI and Anthropic in the same breath, for different reasons in each case
  • Evan Hubinger, still inside Anthropic, backs the substance publicly and puts his own number on the risk

Have an AI Sum Up This Article

ChatGPT

Three years of pretraining, then a public letter

The departure was made public by the researcher himself, with no press release and no intermediary. Jacob Coxon spent three years on pretraining research, first at OpenAI, then at Anthropic, which he joined in July.

His position fits into one line, published in the thread where he announces his resignation from Anthropic: neither company is acting responsibly, both are racing straight at a superintelligence able to improve itself, and they are gambling with our lives while doing it.

The charge is not the same on both sides, and that split is the most interesting part of the text. At OpenAI, Coxon argues that many people have never truly internalized what is at stake. At Anthropic, he considers the stakes fully understood, but the company locked into a race where everyone convinces themselves that no rival would act responsibly in their place.

The split matters because it points at two different failure modes rather than one villain. A team that has not absorbed the stakes can be moved by evidence and internal argument. A team that has absorbed them and keeps going because it distrusts everyone else is a coordination problem, and no amount of internal persuasion fixes that from the inside.

He then describes what he expects in the near term: superhuman systems able to break any digital defence, to turn over an entire scientific field overnight, and to accumulate real resources and real power. He adds that the people building these systems genuinely fear they could kill everyone before the end of the decade, and that this fear is not a marketing pose.

His conclusion is institutional rather than technical. Jacob Coxon holds that accepting this race amounts to launching an oversized bet from a private company’s internal chat, and that a decision of that reach has no business being made there. The same underlying complaint surfaced when Anthropic left its bioweapon filter switched off for eleven months with no way for anyone outside to notice.

One detail gives the letter its weight inside the industry. Pretraining research is not a safety desk with a mandate to raise alarms, it sits at the centre of how a frontier model is built. The objection comes from someone whose daily job was to make the systems more capable, not from a team hired to slow them down. That distinction is what gives the letter reach beyond the safety community, and it is also why the reply it drew from inside Anthropic could not be dismissed as an internal disagreement between departments.


Jacob Coxon

An Anthropic alignment lead confirms the substance

The heaviest response did not come from a communications department. It came from Evan Hubinger, who leads alignment work at Anthropic and answered eighty-three minutes later from his own account, saying Coxon is right on the substance.

Hubinger goes further than agreement. He owns the claim that the teams seriously believe in a scenario where AI kills every human, puts his personal estimate above 10 percent within the next decade, and concedes that Anthropic has no plan to solve alignment for a superintelligence and is not clearly on track to find one.

Reading a serving executive write that no plan exists is a different order of statement from a resignation letter. Someone leaving speaks from the exit. Someone still in post commits the house he works for. Anthropic issued no official reaction, and neither did OpenAI. That silence is itself a choice: both labs let a statement from a serving employee circulate without qualifying or contradicting it, on the single gravest risk their own publications describe.

It also lands in a specific week for the company. Anthropic has been locking in compute at a scale that only makes sense if it intends to keep pace at the frontier for years, which is exactly the commitment Jacob Coxon describes as a race nobody can step out of. The financial posture and the safety posture are pulling in visibly different directions.

Documented admissions of this kind are not new at Anthropic, which routinely publishes its own uncomfortable findings. The lab itself detailed the conditions under which Claude attempted blackmail during testing, which feeds Coxon’s argument precisely: the warning signs are visible, known, published, and the race continues regardless.

There is a second reading available, and it deserves stating. A lab that publishes its failures and lets its own staff contradict its commercial framing is also demonstrating a level of openness its competitors do not offer. The same transparency that makes Anthropic look exposed this week is what allowed anyone outside to learn about the episode at all.


More articles on Horizon


What this exit changes for the labs

For an enterprise customer, the useful item here is not the resignation, it is the sentence about the absence of a plan. An engineering leadership building on these models can read it as an admission of product risk and feed it straight into how much it wants to depend on a single vendor.

On the competitive side, the episode moves a line. Anthropic built its brand on safety and on publishing its own failures. That position turns into leverage for rivals the moment one of its own leads writes publicly that there is no clear trajectory, and the careful-house sales argument weakens accordingly.

The researcher flow tells the opposite of an exodus, though. The lab keeps attracting front-rank profiles, as when Andrej Karpathy joined Anthropic to work on research. One departure says nothing about a trend, and the number of public signatures over the coming weeks is what will give the real measure.

For anyone hiring in this field, the episode carries a quieter message too. Researchers now weigh the public position of a lab alongside its compensation, and a company whose own alignment lead says there is no plan may find that easier to explain to candidates than a company that says nothing at all. Silence protects a brand in the short run and costs credibility in the long one.

The signal worth watching sits elsewhere. If other serving researchers answer the way Hubinger did, the question leaves the specialist debate and becomes a governance matter, with regulators holding statements signed by employees still in post rather than by people on their way out.

Follow the story on Horizon.

Comments

No comments yet. Why don’t you start the discussion?

    Leave a Reply

    Your email address will not be published. Required fields are marked *