Skip to content
AI Primer
breaking

Anthropic says Claude designed binders for 14 of 15 protein targets

Anthropic says Claude ran computational protein-design campaigns from expert-written protocols, producing binders tested by independent labs. Across 1,320 designs, reported hit rates ranged from 22% to 35%.

5 min read
Anthropic says Claude designed binders for 14 of 15 protein targets
Anthropic says Claude designed binders for 14 of 15 protein targets

TL;DR

  • Claude designed lab-measured binders for 14 of 15 targets with interpretable assays, as Anthropic's campaign announcement reports, after starting across 16 targets.
  • The pooled outcome was 354 binders from 1,320 measured designs, a 26.8% hit rate, according to rohanpaul_ai's campaign summary.
  • Campaign shape moved the result: kimmonismus's results breakdown lists 22.6% for Opus 4.8 and 26.7% for Mythos Preview in multi-target runs, versus 35.1% for Mythos Preview in single-target runs.
  • The experiment establishes high-affinity binders as an early drug-discovery input; Anthropic's drug-development caveat says safety and efficacy require many subsequent stages.

The technical report puts a 16,000-word protocol in every agent's system prompt, with most of it devoted to orchestration and operations rather than biology. The public dataset includes raw sensorgrams and report images from two contract research organizations, alongside the design provenance. Adaptyv Bio says it received sequences anonymously for wet-lab testing.

16 targets, 1,320 measurements

Anthropic gave Claude 16 protein targets, then excluded mature GDF-8 after it aggregated under assay conditions at both labs. The resulting analysis covers 1,320 designs on 15 targets, with at least one binder on 14.

The report separates two operating modes:

  • Multi-target: Opus 4.8 and Mythos Preview each ran a 48-hour campaign against the same 14 targets, using a $50,000 cloud-GPU budget per campaign.
  • Single-target: Mythos Preview ran 24-hour campaigns on all 16 targets, while Opus 4.8 ran three; each single-target campaign had a $10,000 GPU budget.

Each campaign delivered 30 ranked designs for a target. Two CROs, Adaptyv Bio and Twist Bioscience, synthesized and tested the delivered sequences, which the report says had been shuffled without model, campaign, or rank labels.

System prompt and toolchain

The experiment gave Claude a campaign procedure rather than a target-specific epitope, scaffold, or protein sequence. daniel_mac8's technical-report notes highlights that the only mid-campaign human messages were short requests to resume after infrastructure failures.

According to the technical report, the prompt allocated 34.2% to science and tooling, 34.7% to orchestration and verification, and 31.1% to operations. The agent was asked to:

  • Research target biology and available structures.
  • Choose target constructs and epitopes.
  • Build and validate public design and prediction tools from source repositories during the first hour.
  • Remove known-protein lookalikes, duplicates, and problematic sequences before expensive scoring.
  • Screen with one predictor seed, rank with five, then run additional in-silico optimization when warranted.
  • Return 30 ranked sequences and a record of its decisions.

The ranking ensemble used ESMFold2, ESMFold2-Fast, and Protenix v2. The agent could draw from generators including PXDesign, RFdiffusion3, Genie 3, and FreeBindCraft, then use SolubleMPNN or related tools for sequence design.

Humans still chose the targets, supplied the reference corpus and cloud account, placed synthesis orders, and interpreted experimental data. The report says the operator also approved access prompts carrying no scientific content.

Hit rates and ranking

The 26.8% pooled figure combines campaign formats. The technical report reports 26.7% for Mythos Preview and 22.6% for Opus 4.8 in 48-hour multi-target sessions, then 35.1% for Mythos Preview in 24-hour single-target sessions; it cites 10% to 15% as typical for protein-design campaigns.

The highest-ranked design in each campaign bound 49% of the time. On RBX1, Claude produced 28 binders from 90 designs, while an open competition had found nine from 245; the report measured the tightest Claude binder at 3.9 nM and the competition winner at 45 nM on the same plate.

Adaptyv's case study says the designs would have won five of six of its prior protein-design competitions on hit rate and binder tightness. The lab's workflow progressed from digital sequence to DNA, expression, binding measurement, and quality control.

Binding assays

The reported metric is the fraction of designs that bound their intended target. Anthropic describes a protein binder as an early component of drug development, with later work needed to establish whether a molecule is safe and effective in Anthropic's own caveat.

Adaptyv used cell-free expression and SPR/BLI kinetics with the design immobilized; Twist used Fc-fusion expression and capture SPR with a six-point antigen titration, as the dataset card documents. The report also says no experimental structure of any designed complex was determined, so every reported binding pose remains a prediction.

The campaign's confidence scores did not reliably separate a fully failed campaign from a successful one, rohanpaul_ai's report summary notes: unsuccessful targets could receive computational scores similar to successful targets.

Prompts, provenance and raw data

Anthropic has released the system prompts, kickoff messages, computational models, and measurements. Anthropic's release post announced the prompt and data publication alongside the technical report.

The Hugging Face release contains 1,440 de novo miniproteins, 50 to 120 residues long, across 16 targets: 900 designed by Mythos Preview and 540 by Opus 4.8. Per design, it links both vendors' binding calls and kinetics, raw sensorgrams, report images, a final cross-vendor assessment, the model used, and structural predictions.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 6 threads
TL;DR3 posts
16 targets, 1,320 measurements1 post
System prompt and toolchain2 posts
Hit rates and ranking1 post
Binding assays2 posts
Prompts, provenance and raw data1 post
Share on X