AI NewsDev toolsReported

CodeScene says two engineers rewrote 300,000 lines of C in three weeks using coding agents, and practitioners are questioning what the setup actually measures

CodeScene says two engineers used coding agents to rewrite a 300,000-line C codebase, the decompiled source of Street Fighter III 3rd Strike, from a Code Health score of 5.6 to 10.0 in three weeks at a token cost the firm reports as $4,000. Practitioners are pressing on the scope, the metric and the harness before treating the result as general.

AI News

Editorial3 min read

LinkedInX
CodeScene case study hero card for the Street Fighter III refactor, published by CodeScene

Image: CodeScene

Why it mattersA number this large will be quoted at every engineering budget meeting in the next month, so a team building software needs to understand what parts of the setup are unusual before promising the same result on its own code.

Two engineers say they took a 300,000-line C codebase from a CodeScene Code Health score of 5.6 to a perfect 10.0 in three weeks of part-time work, at a token cost CodeScene reports as $4,000. The codebase is the decompiled source of Capcom's Street Fighter III 3rd Strike. CodeScene, whose tool graded the result, published the case study on 2026-09-24, and InfoQ's coverage on 2026-09-30 collected the reactions of named practitioners.

CodeScene founder Adam Tornhill writes that after three decades on large-scale systems this is the first time he has seen what he calls superhuman AI performance at scale. Daniel Webb, CTO at NeoSee and one of the two engineers on the work, replied on the discussion thread that the changes were merged to a fork's main through 54 pull requests. The team settled on Claude Opus for most of the run, reporting that Claude Code with Opus captured and documented the emerging patterns more reliably than Codex with Sol, and that files often plateaued when smaller models ran the same task.

Two things the setup does that most teams cannot

The first is a quality signal the agents could act on: CodeScene's Code Health score, computed on every commit as the target for each iteration. The second is a correctness oracle: a replay-trace harness that compared a rollback state hash frame by frame against the original binary. A decompiled fighting game is deterministic, so an exact-match comparison is possible, and Tornhill writes that automated tests and equivalence checks are absolutely essential safeguards. The safeguard is exactly what unhealthy codebases tend to lack. Webb told InfoQ the team considered a Gov.UK marine licensing codebase he had worked on, but it was, in his words, too healthy for the research that follows, so they picked the game partly because they play it.

The playbook is the more novel result

Rather than apply a fixed catalogue, the agents accumulated 22 recipes and 82 supporting notes over the three weeks, and kept a record of failed attempts. Familiar transformations appear, including Extract Function and Guard Clauses. Others are specific to this codebase, such as Shared Index Range for repeated loops that differ only in their bounds and Action Parameter for duplicated control structures that differ mainly in which function they invoke.

The critique from practitioners

Denis Baltor argued to InfoQ that DRY concerns duplication of knowledge, not identical lines of code, and that at least two of the recipes may be collapsing structural duplication with duplicated intent. Asko Nõmm noted that Claude Code and Codex are each tuned to their own models, so it is unclear how much of the result measures the model and how much measures the harness, and that architecture remains unmeasured, so code can look healthy while fundamental problems show up later. Konrad Otrębski asked whether this was open source rather than production code earning money. Marc Bouvier asked whether non-functional behavior improved, given that framerate, memory use and input latency matter in a game; Webb said a performance specialist was being brought in, and that the harness ran as a pre-commit hook, so some frame-timing regressions were fixed without being observed.

Two figures that will get quoted from the case study belong to future work rather than this one: CodeScene projects a 70 percent reduction in AI-induced defects and 45 percent less token waste when future features are built on the uplifted code, both extrapolated from earlier CodeScene research and not measured here. The token cost and the three weeks were measured, on a codebase whose verification method most enterprise systems do not have.

Source

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX