AI NewsModels & agentsAnnouncement

Ninjahawk's livenerf runs a daily benchmark to catch Claude Opus 5.5 getting worse after launch

Livenerf runs the same frozen panel of questions against Claude Opus 5.5 every day for thirty days, and the first call on whether the model has got worse lands on 24 October 2026.

AI News

Editorial2 min read

LinkedInX
ninjahawk/livenerf GitHub repository card

Image: GitHub

Why it mattersA team that pays for a frontier model has had no clean way to answer the "is it getting worse" question, and a pre-registered daily benchmark converts the usual argument from opinion into a readable chart.

A team that pays for a frontier model has had no clean way to answer the question that keeps coming back on forums: has it got worse since launch? Livenerf, a repository published by GitHub user ninjahawk on 22 September 2026, is a pre-registered daily benchmark that aims to answer that question for Claude Opus 5.5. The project collected 899 stars in its first eight days.

The project runs on launch-day samples that were already locked, so a baseline exists before anyone argues about drift. Claude Opus 5.5 was released on 22 September 2026, which is also the repository's day zero. Livenerf's first ten days are the baseline window, and the first "has the model drifted" readout is scheduled for around 24 October 2026.

The method, and what the author claims about it

Livenerf is built on Inspect, the UK AI Security Institute's open-source eval framework. The benchmark freezes everything the author can hold still (prompts, graders, raw logs and the Claude Code command-line version) and measures drift statistically across thousands of samples. The statistical approach follows Anthropic's own 2024 paper on error bars for evals.

The question panel was chosen by sampling 2,336 questions from GPQA Diamond, MMLU-Pro, competition-math and AIME 2025 to 2026, with four samples each. Opus 5.5 gets about 93 percent right on the first try, and 97 percent of questions are always right or always wrong. The 78 that are sometimes right form the panel. One run a day of the whole panel detects an accuracy change of about 7.5 points per ten-day window, according to the author's own power calculation.

What validation found

Before the series began, ninjahawk ran a positive-control experiment to show the setup can see a real change. Running Opus 5.5 at low reasoning effort cut output tokens by 62 percent and dropped accuracy by 8.3 plus or minus 4.5 points. At medium effort, tokens fell 26 percent and accuracy by 4.2 plus or minus 3.9 points.

The limit is also stated plainly in the repository. Swapping Opus 5 in place of Opus 5.5 was not distinguishable at 99 percent confidence over a validation sample (minus 3.8 plus or minus 6.3 points of accuracy, minus 23 percent of tokens). The author says the instrument cannot detect a same-family model swap of that size in a validation-sized sample.

The author also notes an audit of the 78 panel questions found eight answer keys that look wrong and 30 ambiguous questions, which is what the method would predict for questions a strong model only sometimes gets right. Nothing was dropped, but a pre-registered sensitivity analysis will rerun the result without them.

For a team buying a frontier model under a subscription with no official degradation statement, a public, dated chart is the first thing that can turn the usual thread of complaints into a measurement with error bars on it. The caution is that these are early days: the project has one baseline window of data, and the strongest result livenerf can report is a 7.5-point accuracy shift on one specific model, read against the author's own frozen panel.

Source

ninjahawk/livenerf on GitHub, repository created 22 September 2026, star count read 1 October 2026.

SourceGitHub

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX