Claude Designed Proteins That Worked in the Lab

Claude designed proteins shown as a giant molecular key unlocking a vault in a workshop

Claude designed protein binders that hit fourteen of fifteen targets, with a 26.8 percent success rate measured in a wet lab. Anthropic handed the checking to two outside vendors instead of running it in house. The full recipe went public, including a system prompt that runs to sixteen thousand words.

Key Takeaways

  • Of 1,320 designs synthesised and tested, 354 bound to their target, a 26.8 percent rate against the 10 to 15 percent a normal campaign returns.
  • On the RBX1 target, the best design binds roughly ten times tighter than the winner of a public design contest run by Adaptyv Bio.
  • Anthropic states plainly that no human control campaign ran alongside it and that none of the binders were structurally validated.

Have an AI Sum Up This Article

ChatGPT

Fourteen targets out of fifteen, checked in a wet lab

Claude designed these binders by driving a chain of computational biology tools on its own, from backbone generation through to candidate filtering. Two models did the work, Mythos Preview and Opus 4.8. Fourteen of the fifteen protein targets ended up with at least one confirmed binder.

The headline number is 26.8 percent, meaning 354 working designs out of 1,320 made. In multi-target mode Mythos Preview lands at 26.7 percent and Opus 4.8 at 22.6 percent. In single-target mode, with 2.8 times the compute budget, Mythos Preview climbs to 35.1 percent. The designs the model itself ranked highest bound 49 percent of the time.

A working binder campaign today lands somewhere between 10 and 15 percent. The gap here comes from volume as much as from rate, since the final screen ran across more than a thousand sequences. Synthesis and measurement went to Adaptyv Bio and Twist Bioscience, neither of which took part in the design stage.

RBX1 is where the spread gets uncomfortable for the field. Claude hit 40 percent on that target while human entrants in a design contest topped out at 3.7 percent, and the vendor that built and measured the sequences walks through its testing protocol. Its strongest candidate binds at 3.9 nanomolar against the contest winner’s 45.

On TNFα, Opus 4.8 produced binders that work across the human, monkey and mouse versions of the protein. Cross-reactivity of that kind is what makes preclinical testing possible at all. Across the whole run, 130 designs out of 233 also bound the mouse variant, which extends the direction Anthropic took when it pointed Claude at neglected diseases alongside pharmaceutical partners.

Two targets held out. The synthetic beta barrel BBF-14 returned nothing conclusive, and neither did maltose-binding protein, despite ninety designs going through the test. Anthropic publishes both failures at the same level of detail as the wins.


Claude Designed

Fifty thousand dollars a campaign, no expert in the loop

Cost is the figure most likely to travel. A multi-target campaign runs to roughly 50,000 dollars, a single-target one to 10,000, with up to 12,500 H100 GPU hours behind the first. That economics sits directly on top of the platform unveiled in early July, when Anthropic opened a science workbench aimed at laboratories.

No new biology model enters the picture. Every building block is already public, from RFdiffusion3 and PXDesign for backbones to SolubleMPNN for amino acid chains and ESMFold2 and Protenix v2 for prediction. The model conducts more than a dozen tools spread across twenty-four workflows, with nobody stepping in mid-run. Compute ran on the cloud provider Modal.

Structural variety is the quieter result inside the numbers. Fifteen of the binders across six different targets carry beta sheets, a shape that generative pipelines have historically struggled to produce reliably. High-affinity binders came back on at least six targets, and several designs matched or beat the best previously published results for the same protein. That spread matters more than any single hit rate, because it suggests the chain is not narrowly tuned to one easy structural family.

A second strand drew less attention. Opus 5 read raw NMR data in twenty-three minutes and LC-MS data in nineteen, working from proprietary instrument formats. Its hydrogen counts per peak came within 0.08 of the lab’s, and it put purity at 96.4 percent against the lab’s 96.33. This is the same model that we flagged as matching the top of the market at half the price.

For a research team the practical shift shows up in who does what. When Claude designed and sequenced the steps itself, orchestration expertise lost value while judgement on framing the target and reading the output gained it. The bottleneck moves back to the bench, because someone still has to make the thing and measure it.

Whether that orchestration holds up is the open question. Two thirds of the system prompt goes to sequencing and budgeting, which says a great deal about the effort needed to keep an agent on a long chain. We documented the difficulty six days ago, when three Claude agents on one project got in each other’s way.


More articles on Horizon


The sixteen-thousand-word prompt is public

Anthropic put the prompts, the design datasets and the experimental measurements online next to the technical report. The lab laid out the method and results in its research write-up on accelerating protein design. Nothing in the described chain is reserved for its own teams.

That is the part rivals will feel. Any lab with a frontier general model, a compute provider and a synthesis vendor can rerun the exercise without a special licence. The entry barrier is no longer access to design tools, it is the ability to write the sixteen thousand words of instructions that hold them together.

The caveats Anthropic attaches deserve reading before anyone extrapolates. No parallel human campaign served as a control, which rules out any claim of beating equipped specialists. Only 354 designs went through experimental testing, no structure was validated, and four of the six contest targets appeared in the material the model had read.

The safety file gets heavier too. A general model that can run a protein design chain end to end becomes a policy object as much as a research tool, and the company has already stumbled on that ground when a filter meant to catch biological risk sat switched off for eleven months.

Replication decides what this is worth. Until an outside group runs the published protocol against a campaign led by specialists, 26.8 percent stays an in-house result checked by third parties rather than a benchmark. Releasing the data is what makes that comparison possible now.

Follow the story on Horizon.

Comments

No comments yet. Why don’t you start the discussion?

    Leave a Reply

    Your email address will not be published. Required fields are marked *