Jake Schincariol released arena-skill, a Claude Code plugin that runs 100 copies of Claude on the same task and picks the winner
Jake Schincariol released arena-skill under MIT on 27 September 2026, a Claude Code plugin that runs 100 copies of Claude on the same task with different strategy cards, has them critique each other across a seven-round bracket, and returns the one answer that survives.
Image: GitHub
Why it mattersIt is a workflow a working developer can run on a hard prompt tonight, and the author's README puts the cost of a full tournament at 595 sub-agent calls so a team can decide whether a better answer is worth the token spend.
Claude's first answer is often not its best, and the usual fix is to type the question again differently and hope. Jake Schincariol published a different fix on 27 September 2026: arena-skill, a Claude Code plugin that runs 100 copies of Claude on the same task, has them critique each other, and returns the one answer that survives the bracket. The repository picked up 116 GitHub stars in its first four days, and is MIT licensed so a reader can see what each round actually asks the model.
How the tournament runs
The README says the plugin spawns up to 100 sub-agents, each with the same task text and one card drawn from a deck of 2,160. A card is one of 15 reasoning modes, 12 workflows and 12 strategies combined. Named modes in the README include first principles, inversion, analogy, adversarial, constraint-first, worked-example, Socratic, contrarian and systems thinking; named workflows include draft, critique, rewrite; outline-first; test-first; research then synthesise; three drafts, pick one. Pairs then run an attack phase, with each side listing up to seven flaws in the other's answer tagged FATAL, MAJOR or MINOR. Each agent rebuts and revises. A separate Claude agent judges both revised answers and the lower score is cut.
What it costs to run
The author quotes the token cost in the README: a 100-agent bracket takes 595 sub-agent calls over seven rounds (100 to 50 to 25 to 13 to 7 to 4 to 2 to 1), and a --quick flag drops that to 91 calls across 16 agents. The judging rubric is 30 points correctness, 25 completeness against the task, 20 robustness to the attacks raised, 15 specificity and 10 clarity, totalling 100. Flags cover a custom agent count, a seed for a reproducible bracket, and a wave size for how many sub-agents run at once.
Where the traction sits
116 stars in four days works out to about 29 a day, enough to put the repository on the GitHub trending page. The same author released a Claude skill for writing LinkedIn posts in September that reached 655 stars, so this is a known maker in the Claude Code skill world. Arena-skill is a workflow wrapper that sits around any job-specific skill a team already uses, for the tasks where a single answer is not good enough.
The honest limit of a tool like this is that 595 model calls is a real token bill, so a team will reach for it when the prompt is high-value and not on everyday work. The judge is itself a Claude agent, so the scoring is only as good as the rubric and the opponents' attacks, and a weak attack phase will promote a confident wrong answer. The seed and the strategy deck make a run reproducible, which is the smaller part of the honest case and the one that lets a team compare different prompts under the same bracket.
Source
- arena-skill repository, MIT, Jake Schincariol, 27 September 2026
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.