Livenerf benchmark tracks whether Claude Opus 5.5 gets worse after launch

Livenerf benchmark tracks whether Claude Opus 5.5 gets worse after launch

Livenerf is a GitHub repository that sets out to answer one question: does a frontier model get worse after it ships? The README notes that for months there have been reports that Anthropic "nerfs" models some days or weeks after release. That could mean quantization, a smaller model behind the same name, lower effort, or routing changes. It could also mean nothing happened and people are pattern-matching on noise. Nobody has had a clean day-0 baseline to check against, so arguments have come down to vibes versus vibes. Claude Opus 5.5 came out on 2026-09-22, so the author started the clock right after launch.

The models cannot be made deterministic: sampling parameters are gone and thinking cannot be turned off. So livenerf freezes everything else (prompts, a pinned Claude Code CLI, exact graders, raw logs kept forever) and measures drift statistically. It is built on Inspect, the UK AI Security Institute's open-source eval framework, and the statistics follow Anthropic's own "Adding Error Bars to Evals". Version 0 runs on a Claude Max subscription through headless Claude Code (claude -p), with no API key. Scoring is exact-match, with no LLM judge, because a judge could drift too.

The series is live. Day 1 was 2026-09-24 22:10 UTC, about 2.5 days after launch. It runs once a day for 30 days: days 1 to 10 are the baseline, then come two 10-day windows, so the first possible call is around 2026-10-24. The first Results row lands after day 20. As of 2026-09-29, 6 of 30 days were collected (baseline 6 of 10), none missed. All six ran the full 90 samples on the same harness hash and pinned CLI version 2.1.280. Day 5 ran with the budget guard overridden once, noted in a deviations log.

The question panel was built under a pre-registered protocol (v2). The author screened 2,336 questions from GPQA Diamond, MMLU-Pro, competition math and AIME 2025 to 2026, with 4 samples each. Opus 5.5 gets about 93% right on the first try, and 97% of the questions were always right or always wrong. Only 78 questions are sometimes right, and those form the panel, since a question the model always gets right cannot show a drop. Selection bias was measured: on fresh samples the panel's pass rate rose from 54.7% to 62.0%, so the power calculation uses the fresh rates. With one whole-panel run a day, the design can detect an accuracy change of about 7.5 points per 10-day window, at a cost of about 3.6% of the weekly plan.

Validation passed its pre-registered criterion. Lower effort showed up much more clearly in tokens than in accuracy: effort low cut output tokens by 62% and accuracy by 8.3 ± 4.5 points; effort medium cut tokens by 26% and accuracy by 4.2 ± 3.9 points. The limit is that swapping in Opus 5 was not distinguishable from Opus 5.5 at 99% (accuracy down 3.8 ± 6.3 points, tokens down 23%). The README says the instrument cannot detect a same-family model swap of that size in a validation's worth of samples. A 10-day window has about 2.5 times as many samples, but that has not been shown to be enough.

The questions themselves are imperfect. A report-only audit of the 78 questions, plus 2 later excluded, found 8 answer keys that look wrong and 30 ambiguous questions. Nothing was dropped; a pre-registered sensitivity analysis reruns the result without them. On the serving path, the safety classifier sometimes answers with Opus 5 or refuses biology and some math questions. Those samples are rejected and counted, and questions it touched are excluded.

The decision rule is fixed in advance. The primary metric is the paired per-item score difference against the baseline, with clustered standard errors, so item difficulty drops out. A change has to clear a 99% interval in both 10-day windows that follow, be at least 3 points, and not show up in the control arm. The control arm runs claude-opus-5 on the panel's GPQA questions through the same harness every day, to separate a harness or platform change from an Opus 5.5 change. The author's secondary signal is output token count per sample: if a model quietly starts thinking less, it may show there before accuracy moves. The main output will be a running 10-day table relative to the launch-week baseline, reporting improvements as loudly as regressions.

Key facts

  • Livenerf is a pre-registered, 30-day daily benchmark of Claude Opus 5.5, started 2026-09-24, about 2.5 days after the 2026-09-22 launch; the first possible call is around 2026-10-24.
  • As of 2026-09-29, 6 of 30 days were collected (all baseline days), each with 90 samples on the same harness hash and pinned CLI 2.1.280. There are no drift results yet.
  • The panel is 78 questions, picked from 2,336 screened (GPQA Diamond, MMLU-Pro, competition math, AIME 2025 to 2026) because the model sometimes gets them right; the design can detect about a 7.5-point accuracy change per 10-day window.
  • Validation found effort low cut output tokens 62% and accuracy 8.3 ± 4.5 points, but a swap to Opus 5 was not distinguishable from Opus 5.5 at 99% (down 3.8 ± 6.3 points).
  • A change counts only if it clears a 99% interval in both 10-day windows, is at least 3 points, and does not appear in the control arm.

Why it matters

Claims that a model gets quietly worse after launch are common, and the README says they have lacked a clean day-0 baseline to check against. Livenerf tries to supply one: frozen prompts, a pinned CLI, exact-match grading and a decision rule committed before the data came in. Whatever it finds, the approach makes the argument testable instead of anecdotal. It also tests for change in either direction and does not assume a mechanism.

Who it affects

Anyone who relies on Claude Opus 5.5 and wonders whether it changes after release, especially users on Claude Max subscriptions, since that is the setup the benchmark measures. It also interests evaluation researchers who want a template for long-running drift monitoring with error bars. Anthropic is the subject of the nerf reports, but the README reports no reaction from the company.

How to use it

The repo is open on GitHub (ninjahawk/livenerf). To follow it, watch the Results table, which gets its first row after day 20. To run your own copy you need Python 3.11+, uv and a logged-in Claude Code install; anything that can run claude -p works, and Linux, macOS and Windows are supported. Pin the CLI version and turn off auto-updates, because a changed harness looks exactly like a changed model. Budgets are set in points of the weekly plan meter, and a daily attempt is skipped if the weekly meter is at or above 75% or the 5-hour meter at or above 60%. Running against the API later means changing only the model string to anthropic/claude-opus-5-5.

How solid is it

The method is carefully specified: built on the UK AI Security Institute's Inspect, statistics following Anthropic's Adding Error Bars to Evals, a public pre-registration committed before the series, locked panel and raw logs kept append-only. The author is also frank about limits. Validation could not separate Opus 5 from Opus 5.5 at 99%, and the claim that 10-day windows will be enough is not yet demonstrated. Only baseline data exists, so the project has said nothing about whether Opus 5.5 has changed. The README does not name its author; the repository sits under the GitHub account ninjahawk.

Risks and caveats

Livenerf measures Opus 5.5 as served through Claude Code on a subscription, which is not the same as the raw API model, though the README says that is what most nerf reports are about. The launch-week baseline is a reference point, not ground truth; launch week could be the worst week because of a new serving stack, capacity strain and launch bugs. The README adds that the 2025 quality incidents turned out to be infrastructure bugs, not deliberate downgrades. The panel has flaws: 8 answer keys that look wrong and 30 ambiguous questions, none dropped. The safety classifier can reroute some answers to Opus 5 or refuse some questions, so affected samples are rejected and counted. Day 5 ran with the budget guard overridden once.

“Nobody has had a clean day-0 baseline to check against, so every argument ends up as vibes versus vibes.”

— livenerf README