What good code review looks like when nobody wrote the code

Ask ten engineers what code review is for and you will get answers that quietly disagree. Some think it is a checkpoint that blocks bad changes. Some think it is where you enforce your opinions about how code should look. Some think it is a chore you do at the end of the day so a teammate can continue. Those disagreements did little harm when a person wrote the code, because the author was doing most of the quality work anyway and review was a useful second check.
That is no longer true. When the implementation is generated, there is no first check. The reviewer is the only check, and a process designed as a backup does not become the main control just because the first check disappeared.
What review is for now
Two of the three traditional jobs remain, and one is gone.
Catching problems early still matters, for the same reason it always did. The cost of a defect grows the longer it stays in the code, a point NIST has documented in its work on the economic cost of software errors, and review remains the cheapest place in the process to catch one.
Building shared understanding matters much more. When a person writes a change, at least one human understands it because they wrote it. When a model writes it, that is no longer true. If the reviewer does not genuinely understand the change, then nobody does, and you have added something to your system that no human can explain. Six months later, when it misbehaves, the understanding does not exist to be recovered. It was never created.
Helping the author improve is gone. You cannot teach a model to be better on the next pull request. What replaces it: the feedback now goes into the specification. When a reviewer keeps finding the same type of gap, the fix is a sentence in the spec that should have been there, more than a comment on the change. Reviewing generated code teaches you about your spec, not about your colleague.
Generated code fails differently
This is the part teams underestimate, and it is the reason a reviewer cannot simply keep doing what they were doing.
Human code fails in human ways. Someone forgets a case, misreads a requirement, gets an index off by one, or runs out of attention at 6pm. Experienced reviewers are trained to notice exactly these, and much of that skill is unconscious. It is pattern recognition built over years of watching the mistakes people typically make.
Generated code does not make those mistakes, and that trained skill does not help with it. Its characteristic failures are different:
It is plausible and wrong in small ways. The code reads well, follows the conventions, and solves a problem close to yours. Close is the dangerous part. A reviewer skimming for obvious errors finds none, because there are none of that kind.
It fills gaps silently. Anywhere the specification left a decision open, something was still decided, and nothing in the change marks where. There is no comment saying "unclear from the spec, assumed the first case wins." The assumption is just there, in the logic, looking exactly like a requirement.
It does more than was asked. Generated code often handles cases nobody asked about, sometimes elaborately. This looks like thoroughness and is frequently the opposite: it is extra code you did not want, behaviour you did not specify, and code you now have to maintain.
It is confident about the main path and incomplete on everything else. This is the most reliable pattern of all. Normal use is usually handled correctly. The timeouts, the retries, the partial failures, the concurrent case: that is where the gaps are, consistently.
The practical consequence is that the reviewer's question has to change. Asking "is this correct?" does not help, because it will look correct. The question is "is this the thing we specified?", checked case by case against the document it was generated from, and most carefully on the unusual paths.
Small changes are now required
If we could change one thing about how a team reviews generated code, it would be size, and the reason is now the opposite of what it used to be.
Small pull requests were always better. What used to keep them small, though, was partly effort: writing a thousand lines took a person a week, so it did not happen casually. AI removed that limit entirely. A thousand-line change now costs the machine nothing to produce, which means nothing is stopping it from arriving except a rule you enforce.
Meanwhile the cost on the other side went up. A reviewer can read a twelve-line human change quickly, because they can rely on the author's judgment for the parts they skim. On a generated change there is no judgment to rely on. Every line is unverified. Reviewing it properly is slower per line, not faster.
So the two costs moved in opposite directions: cheaper to produce, more expensive to check. A team that does not deliberately limit change size will slowly start approving changes without reading them, and a large generated change that gets approved without real reading is the worst result in this whole process. It has the appearance of a control and none of the substance. The DORA research has long found that small batches lower failure rates rather than raising them, and that finding did not stop being true when the author changed.
Automate more, and know what automation cannot do
Everything mechanical should run before a human sees the change: formatting, style, common bug patterns, the full test suite, dependency and licence scanning. This was good practice before and it is close to mandatory now, because the reviewer's attention is the scarcest resource in the process and every minute of it spent on something a tool could check is a minute not spent on behaviour.
Using a model to review generated code belongs in this category. It is a genuinely useful filter and it catches real things. It is not the last step. It shares failure modes with the thing it is checking, it is subject to the same confident-and-wrong pattern, and it cannot be accountable for the outcome. Use it to reduce what reaches the human, never to replace the human.
Because accountability does not change. Every change we release has a named engineer who read it and answers for it. That is the reason the process exists. Nobody can escalate to a model at two in the morning, and a system whose only author was a model has nobody to ask.
What this looks like day to day
Changes arrive small, because a rule keeps them small. The automated checks already pass, so the routine checks are done. A reviewer picks one up and reads it against the specification rather than against their taste, checking the error paths first because that is where the gaps are. They are direct about the code, because there is no author to discourage. When they find a gap that came from an ambiguous spec, the comment goes on the change and the fix goes in the spec, so the same gap does not arrive again next week.
Then someone puts their name on it. That is the part that makes the rest of it mean something.
Good review was always about two things: problems caught while they are small, and more than one person understanding the system. Both of those got harder to achieve and more valuable to have. The practice did not become less important when people stopped typing the code. It is now where nearly all of the engineering judgment happens, which is the same reason we argue for judging engineers on outcomes rather than volume in measure engineers by outcomes not output.
Sources
- NIST, National Institute of Standards and Technology (research on the economic cost of finding and fixing software defects later in the lifecycle): https://www.nist.gov/
- DORA, DevOps Research and Assessment (research on batch size, deployment frequency, and change failure rate): https://dora.dev/research/


