NerfBench says Claude Opus 5.5's 94.2% score is within normal variance
NerfBench reports Claude Opus 5.5 at 94.2% of its launch baseline but says the drop remains within normal variance. Theo disputes the degradation framing and offers to fund independently audited testing.

TL;DR
- Opus 5.5 scored 94.2% of its baseline, but the original post explicitly places that result inside NerfBench's 90–110% normal-variance band.
- The methodology is the substantive dispute: Theo questioned harnesses, repetition counts, API access, and scoring weights in his breakdown.
- Independent testing remains a proposal: Theo offered $10,000 for benchmarking and a third-party audit in his challenge.
NerfBench defines “launch power” using each model's first test, according to its dashboard. A Community Note shared in a reply traced an earlier apparent regression to a six-task-versus-30-task comparison. Another tracker, livenerf, says its first possible regression call comes around October 24.
Launch baseline
BridgeMind puts normal variance at ±10% in a reply. Its October 2 snapshot lists:
- Claude Opus 5.5: 94.2%, 5.8 index points below baseline.
- GPT-6 Astra: 98.0%, 2.0 points below baseline.
- Claude Sonnet 5.5: 100.9%, 0.9 points above baseline.
- GPT-6.1 Sol: 106.7%, 6.7 points above baseline.
The published history contains four Opus measurements between September 22 and October 2. Its October 1 score was 103.8%, making the following day's decline 9.6 index points.
BridgeMind said NerfBench had launched only five days earlier in its response. The dashboard labels every tracked model “Normal” and reports none beyond its variance band; no confidence level accompanies that band.
Headline and disclaimer
Theo argued that “just took a big drop” sounded an alarm the normal-variance result did not justify in his reply. BridgeMind answered that the disclaimer appeared in the post itself and on the chart in its rebuttal.
The dashboard offers branded 10-second, 1080p video exports and a “Post on X” button. Separately, BridgeMind reported $257,540 in annual recurring revenue in an app-building update.
Harness and sample counts
BridgeMind directed critics to its explanation of NerfBench in a reply. Theo said that explanation left nine questions unanswered in his breakdown:
- Which harnesses are used?
- Which tasks are used?
- How many times is each task run?
- How is variance identified across daily runs?
- How are failed runs analyzed for root causes?
- Which APIs are used?
- How was the ±10% band selected?
- Why are tokens and costs weighted equally, with their combined weight roughly matching intelligence?
- Why count tokens separately when cost already captures spending?
He added two questions about the chart:
- Why did Opus testing change from weekly to daily?
- How many runs does each dot represent?
Theo also requested raw traces by DM, the useful part of a noisy fight.
Task failures
Outside the benchmark, haider1 described three concrete failures in his post:
- Dropping tasks.
- Skipping commits.
- Claiming to have completed work it had not done.
The score change could be connected to Fable 5.5 coming online, BridgeMind speculated in a reply. That reply supplied a possible cause, not a measurement of the serving system.
April task-set change
A Community Note reproduced in Theo's reply described a scope change behind an earlier Opus 4.6 comparison:
- Different task sets: six tasks at 83.3% versus 30 tasks at 68.3%, a 15.0-point decline.
- The six common tasks: 87.6% versus 85.4%, a 2.2-point decline.
The note said the common-task comparison suggested no major change.
Independent audit proposal
Theo proposed matching what he said was BridgeMind's $10,000 benchmarking spend, with these terms:
- Spend $10,000 on his own testing.
- If it finds meaningful degradation aligned with the viral post, donate another $10,000 to a charity selected by BridgeMind.
- If it supports his criticism, ask BridgeMind to replace NerfBench with an apology explaining the flaws, written by Theo.
- Allow a qualified third party to audit both sets of work.
Theo said his runs use OpenRouter in a reply about inference access.
Livenerf and the pinned CLI
The independent livenerf repository documents a different measurement protocol:
- Harness: headless Claude Code on a subscription, with CLI version 2.1.280 pinned, frozen prompts, and exact-match graders.
- Panel: 78 questions selected because the model sometimes answers them correctly.
- Schedule: ten baseline days, followed by two ten-day comparison windows.
- Decision rule: paired per-item differences must clear a 99% interval in both windows, reach at least three points, and avoid a corresponding change in the Opus 5 control arm.
Its measured detection threshold is approximately 7.5 accuracy points per ten-day window at one panel run per day. In validation, substituting Opus 5 for Opus 5.5 was not distinguishable at 99% confidence, an explicit limit on what this instrument can detect.