Engineering

What good code review looks like when nobody wrote the code

Editorial · Reveneau · July 19, 2026 · Updated August 19, 2026

What good code review looks like when nobody wrote the code

Ask ten engineers what code review is for and you will get answers that quietly disagree. Some think it is a checkpoint that blocks bad changes. Some think it is where you enforce your opinions about how code should look. Some think it is a chore you do at the end of the day so a teammate can continue. Those disagreements did little harm when a person wrote the code, because the author was doing most of the quality work anyway and review was a useful second check.

That is no longer true. When the implementation is generated, there is no first check. The reviewer is the only check, and a process designed as a backup does not become the main control just because the first check disappeared.

What review is for now

Two of the three traditional jobs remain, and one is gone.

Catching problems early still matters, for the same reason it always did. The cost of a defect grows the longer it stays in the code, a point NIST has documented in its work on the economic cost of software errors, and review remains the cheapest place in the process to catch one.

Building shared understanding matters much more. When a person writes a change, at least one human understands it because they wrote it. When a model writes it, that is no longer true. If the reviewer does not genuinely understand the change, then nobody does, and you have added something to your system that no human can explain. Six months later, when it misbehaves, the understanding does not exist to be recovered. It was never created.

Helping the author improve is gone. You cannot teach a model to be better on the next pull request. What replaces it: the feedback now goes into the specification. When a reviewer keeps finding the same type of gap, the fix is a sentence in the spec that should have been there, more than a comment on the change. Reviewing generated code teaches you about your spec, not about your colleague.

Generated code fails differently

This is the part teams underestimate, and it is the reason a reviewer cannot simply keep doing what they were doing.

Human code fails in human ways. Someone forgets a case, misreads a requirement, gets an index off by one, or runs out of attention at 6pm. Experienced reviewers are trained to notice exactly these, and much of that skill is unconscious. It is pattern recognition built over years of watching the mistakes people typically make.

Generated code does not make those mistakes, and that trained skill does not help with it. Its characteristic failures are different:

It is plausible and wrong in small ways. The code reads well, follows the conventions, and solves a problem close to yours. Close is the dangerous part. A reviewer skimming for obvious errors finds none, because there are none of that kind.

It fills gaps silently. Anywhere the specification left a decision open, something was still decided, and nothing in the change marks where. There is no comment saying "unclear from the spec, assumed the first case wins." The assumption is just there, in the logic, looking exactly like a requirement.

It does more than was asked. Generated code often handles cases nobody asked about, sometimes elaborately. This looks like thoroughness and is frequently the opposite: it is extra code you did not want, behaviour you did not specify, and code you now have to maintain.

It is confident about the main path and incomplete on everything else. This is the most reliable pattern of all. Normal use is usually handled correctly. The timeouts, the retries, the partial failures, the concurrent case: that is where the gaps are, consistently.

The practical consequence is that the reviewer's question has to change. Asking "is this correct?" does not help, because it will look correct. The question is "is this the thing we specified?", checked case by case against the document it was generated from, and most carefully on the unusual paths.

Small changes are now required

If we could change one thing about how a team reviews generated code, it would be size, and the reason is now the opposite of what it used to be.

Small pull requests were always better. What used to keep them small, though, was partly effort: writing a thousand lines took a person a week, so it did not happen casually. AI removed that limit entirely. A thousand-line change now costs the machine nothing to produce, which means nothing is stopping it from arriving except a rule you enforce.

Meanwhile the cost on the other side went up. A reviewer can read a twelve-line human change quickly, because they can rely on the author's judgment for the parts they skim. On a generated change there is no judgment to rely on. Every line is unverified. Reviewing it properly is slower per line, not faster.

So the two costs moved in opposite directions: cheaper to produce, more expensive to check. A team that does not deliberately limit change size will slowly start approving changes without reading them, and a large generated change that gets approved without real reading is the worst result in this whole process. It has the appearance of a control and none of the substance. The DORA research has long found that small batches lower failure rates rather than raising them, and that finding did not stop being true when the author changed.

Automate more, and know what automation cannot do

Everything mechanical should run before a human sees the change: formatting, style, common bug patterns, the full test suite, dependency and licence scanning. This was good practice before and it is close to mandatory now, because the reviewer's attention is the scarcest resource in the process and every minute of it spent on something a tool could check is a minute not spent on behaviour.

Using a model to review generated code belongs in this category. It is a genuinely useful filter and it catches real things. It is not the last step. It shares failure modes with the thing it is checking, it is subject to the same confident-and-wrong pattern, and it cannot be accountable for the outcome. Use it to reduce what reaches the human, never to replace the human.

Because accountability does not change. Every change we release has a named engineer who read it and answers for it. That is the reason the process exists. Nobody can escalate to a model at two in the morning, and a system whose only author was a model has nobody to ask.

What this looks like day to day

Changes arrive small, because a rule keeps them small. The automated checks already pass, so the routine checks are done. A reviewer picks one up and reads it against the specification rather than against their taste, checking the error paths first because that is where the gaps are. They are direct about the code, because there is no author to discourage. When they find a gap that came from an ambiguous spec, the comment goes on the change and the fix goes in the spec, so the same gap does not arrive again next week.

Then someone puts their name on it. That is the part that makes the rest of it mean something.

Good review was always about two things: problems caught while they are small, and more than one person understanding the system. Both of those got harder to achieve and more valuable to have. The practice did not become less important when people stopped typing the code. It is now where nearly all of the engineering judgment happens, which is the same reason we argue for judging engineers on outcomes rather than volume in measure engineers by outcomes not output.

Sources

  • NIST, National Institute of Standards and Technology (research on the economic cost of finding and fixing software defects later in the lifecycle): https://www.nist.gov/
  • DORA, DevOps Research and Assessment (research on batch size, deployment frequency, and change failure rate): https://dora.dev/research/

Common questions

What is the purpose of code review when AI writes the code?

To confirm the change does what was specified, and to make sure at least one person understands it. The old third purpose, helping the author improve, no longer applies, because you cannot teach a model to be better next time. The first two purposes get more important.

How is reviewing AI-generated code different from reviewing human code?

The typical failures are different. Human code fails in human ways: a forgotten case, an off-by-one, a misread requirement. Generated code fails by being plausible and wrong in small ways, by silently inventing behaviour to fill a gap in the spec, by solving a similar problem instead of the one you asked for, or by adding handling for cases nobody wanted. A reviewer scanning for human mistakes will miss all of these.

What should a reviewer check first in generated code?

Whether it does what the specification said, case by case. Not whether it looks correct, which generated code almost always does, but whether the behaviour matches what was asked for, particularly on the error paths and unusual cases. Correctness of style tells you nothing here.

Does code review get less important when AI writes the code?

The opposite. With a human author there are two independent chances to catch a problem: the person writing it, who has context and doubts, and the person reviewing it. Generated code removes the first one. The reviewer is the only human who will read that change before it reaches production.

Should AI-generated pull requests be small?

More than ever. AI makes large changes cheap to produce, which removes the natural limit that used to keep pull requests small. A thousand-line change costs the machine nothing and costs the reviewer hours of work, and a large change that gets approved without real reading is worse than no review at all.

Who is accountable for AI-generated code?

The named engineer who approved it. Nobody can escalate to a model, so the person who read the change and said it was ready has to be accountable. We treat that as the whole point of the review stage rather than a formality at the end of it.

Can AI review AI-generated code?

It is useful as another automated check, in the same category as a linter or a test suite, and it catches real things. It cannot be the last step, because it shares the failure modes of the thing it is reviewing and it cannot be accountable for the result. Use it to filter, not to approve.

What should still be automated before a human reviews?

Everything mechanical should run before a human sees the change: formatting, style, common bug patterns, dependency and licence scanning, and the full test suite. This mattered before generated code and it matters more now, because the reviewer's attention is the scarcest thing in the process, and every minute spent on something a tool could check is a minute not spent judging whether the code actually does what was specified.

Does tone still matter in code review?

Between people, yes, exactly as much as it always did. On a generated change there is no author to discourage, which lets the reviewer be direct about the code itself. The conversation between people turns to a different question: whether the specification was right in the first place.

How long should reviewing generated code take?

Longer per line than reviewing a colleague's work, which is why the change has to be smaller. If a reviewer is approving generated code faster than they would approve a teammate's, they are judging by how correct it looks rather than checking what it does.