first-pass is a Claude Code plugin that asks ten questions before code and reviews the change in a fresh context after
first-pass is a new MIT-licensed Claude Code plugin that asks ten questions about the code around a change before writing any of it, and reviews the diff in a fresh context with a separate agent before calling anything done. The repository has reached 93 stars in four days.
Image: GitHub
Why it mattersA team using Claude Code can now stop reviewing an agent's work in the same session that wrote it, and can catch the class of bug the author's audit found most: another code path that touches the same data was not updated.
A developer running Claude Code past the "make it work" stage learns quickly that "be 100% sure" changes how sure the agent's answers sound but not what gets checked. first-pass, a new Claude Code plugin published by Joe Tawil, replaces that phrase with ten named checks the agent has to answer against the real code before writing any of it, and a separate agent that reviews the diff in a fresh context afterwards. The repository has gained 93 GitHub stars since 25 September.
Where the ten questions came from
Tawil writes in the README that he built a product feature by feature with Claude Code, then ran a full audit and found 318 issues, 25 of them high severity. He sorted the 318 by cause on one private codebase, by hand, one cause per issue. Only 45 were mistakes in the lines being written. The rest were in the code around those lines: 54 where another code path using the same data was not updated, 54 that ran twice or halfway or at the same time as another call, 45 where an outside service failed and the error was hidden. Time and units, scale limits, endings such as cancel or expire, and UI text that promised what the code did not deliver each contributed their own 15 to 26 issues.
first-pass turns those categories into a premortem skill: before the agent writes code, it answers what happens if the change runs twice, stops halfway, or an outside service times out; which other code uses the same data; what the user sees when it fails. Each answer is a file:line, a test, or "Not handled, because ___" for the developer to accept.
What the "breaker" agent does
A model reviewing its own work in the same context that wrote it is poor at catching its own mistakes, Tawil cites Huang and colleagues, 2024. The plugin's breaker agent runs in a fresh context and reads the diff along with every other code path that touches the same data. "Done" needs a test that failed before the change, the breaker's review, and CI's own checks passing in a clean checkout; anything less is reported as built, not done. On a bug, fix-the-class reproduces the bug, searches for the same pattern elsewhere, and adds the check that stops it coming back.
first-pass is MIT-licensed. It is installed as a Claude Code plugin from a marketplace add and adds session hooks that flag a "done" with no evidence behind it, run each repository's own hooks from a shared parent folder, and note what has drifted since the setup command was last run. Cursor, Codex and other agents get the same rules and skills, but the hooks are Claude Code only.
The audit that opens the README is one product, one developer, sorted by hand. That is not a benchmark and Tawil says so plainly. The number worth keeping is that 273 of the 318 issues were somewhere the agent had not been asked to look. A plugin that names those places by category and asks about them before the code is written is at least addressing the class of bug the audit found most often. Whether these are the ten questions to ask will show after teams have run it on their own codebases. There is a real cost as well, since the premortem uses tokens on every task, and Tawil warns in the README that a run will spend a large share of a usage plan.
Source
Primary source: joetawil7/first-pass on GitHub.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.