AI NewsProductivityReported

CodeScene had agents refactor 300,000 lines of Street Fighter III C code in three weeks for about $4,000 in tokens, and the game still runs frame for frame

CodeScene published a case study on 24 September 2026 in which coding agents rewrote 252,055 lines of C across 2,903 commits over three weeks, guided by a deterministic quality score and verified frame by frame against a game replay. InfoQ reported on the work on 30 September, and practitioners have questioned what one game codebase proves.

AI News

Editorial3 min read

LinkedInX
Before and after Code Health visualisations of the Street Fighter III codebase, showing the file tree turning from red to green

Image: CodeScene

Why it mattersThe bar for agents on a real codebase is no longer "will it write something that compiles", it is "will it change 250,000 lines without changing what the program does", and the answer is now yes when the loop has a deterministic score and a replay harness sitting next to the model.

A large refactor of a legacy codebase used to be a 12 to 18 month project with a high risk of quietly breaking something the tests did not cover. CodeScene, a Swedish company that sells code-analysis tools, published a case study on 24 September 2026 in which coding agents rewrote 252,055 lines of C over three weeks, on a 300,000 line codebase, for about $4,000 in tokens. InfoQ, a developer news site, covered the report on 30 September, and practitioners have started asking what one game codebase actually proves.

What was measured

The codebase is Street Fighter III: 3rd Strike, from an open source decompilation project. CodeScene says the agents produced 2,903 commits across 726 files. The company's Code Health score moved from 5.6 to 10.0, the top of that scale. Adam Tornhill, CodeScene's founder and CTO, credited execution to Daniel Webb and Dr. Markus Borg, and describes the result as the first time he has seen what he calls superhuman AI performance at scale.

The team tried several models and settled on Claude Opus for the main work. CodeScene says smaller models reached a local plateau and stopped finding improvements.

The two feedback loops the agents could not fake

Two pieces of the loop are the whole story here, because either one alone is not enough.

The first is CodeScene's own CodeHealth MCP server, which gave the agent a deterministic quality score after every change. A model that talks to a scorer that always returns the same number for the same code cannot argue its way to a passing grade. It has to actually change the code.

The second is a replay-trace harness that compared rollback state hashes frame by frame before and after each change. That answered the harder question of whether the refactor changed the program's behaviour. If the agent broke the game, the hashes diverged and the change was rejected.

Along the way the team accumulated 22 refactoring recipes with 82 supporting notes, including patterns specific to this codebase such as Shared Index Range and Uniform Step Table.

What practitioners are asking

InfoQ quotes several people who read the report. Paolo Perrone says the replay-trace check sets a higher bar than a typical test suite. Konrad Otrębski questions whether an open source decompilation of a game speaks to a production system a company would actually ship. Asko Nõmm points out that tuning to a specific model leaves non-functional behaviour, meaning speed, memory and how the code fails, unmeasured. Denis Baltor cautions that algorithmic similarity is not the same as semantic intent when the agent applies a rule like "do not repeat yourself".

CodeScene mentions further work with Lund University measuring feature-implementation cost and quality on both the refactored and unrefactored versions, and cites projected reductions of about 70 percent in AI-introduced defects and 45 percent in token waste. Those are projections carried forward from earlier CodeScene research, not results from this case study, and the InfoQ piece flags that.

A team looking at this asks itself two questions before it copies the method. Does the codebase have a deterministic quality score that the model cannot argue with, and does it have a check that catches a behaviour change without a human reading every diff. On this codebase, both answers were yes. On most production codebases, neither answer is yes today.

Source

Primary source: CodeScene, Case Study: Refactoring at Scale with Agents, 24 September 2026. Reporting: Steef-Jan Wiggers, InfoQ, Agents Refactor 300K Lines in Three Weeks, and Practitioners Ask What It Proves, 30 September 2026.

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX