From AI-generated code to production software / Fixing and reviewing
How to review AI-generated code before release
AI-generated code should get a closer read than code from a colleague you trust, not a lighter one, because the failure modes are subtler and speed makes it tempting to skim. Here is a concrete process for reviewing it properly.
Published July 28, 2026. Editorial.
Key takeaways
- Start with automated checks: tests, static analysis, and dependency scanning, before reading a single line yourself.
- Read for assumptions, not just correctness. Ask what the code assumes about its input, caller, and environment.
- Do not trust AI-written tests to validate AI-written code. The same mistake can show up in both.
- Check that the code matches your project's conventions and architecture, not just that it technically works.
Reviewing AI-generated code well is a specific skill, and most engineers were never taught it because it did not exist as a distinct problem a few years ago. Reviewing a colleague's pull request and reviewing an AI's output look similar at first and are different in practice, because the two kinds of author fail in different ways.
Start with what machines can check
Before you read a single line, run the automated checks that exist for exactly this purpose: the test suite, static analysis tools, dependency and secret scanners. Confirm the code compiles, the existing tests still pass, and no known vulnerability was pulled in through a dependency. This step catches a meaningful share of problems for close to no effort, and skipping it means your careful human read is wasted catching mistakes a tool would have found instantly.
If those automated checks are few or weak, that is the more useful thing to fix. A reviewer reading a diff that already satisfies real checks spends their attention on judgment rather than on verifying behaviour by hand, which is the whole argument for eval-driven development.
Read for assumptions, not first impressions
The instinct when reading fast is to ask "does this look reasonable." That question is not strict enough for AI-generated code, because AI output is good at looking reasonable while making a wrong assumption that is hard to see. The better question is: what is this code assuming about its input, its caller, and the environment it runs in, and is that assumption actually true here.
Read the code step by step with that question in mind. If it parses a value, ask what happens when that value is missing, malformed, or hostile. If it calls another part of the system, ask whether it is assuming that call always succeeds. If it handles money, permissions, or user data, ask what happens on the failure path, not just the success path. This is slower than skimming, and it is the entire point of a review: catching the case the author, human or AI, did not think to handle.
Do not let AI check its own work
A common mistake is generating both the implementation and its tests with the same AI tool in the same session. When that happens, the tests tend to validate the implementation's own understanding of the problem rather than the actual requirement, because both came from the same understanding making the same assumptions. High test coverage in that setup can be almost meaningless.
The fix is to keep the original requirement or ticket open while reviewing, and check the code and its tests against that requirement independently, rather than checking the tests against the code. If you can, have a person write or review the test cases separately from the person or tool that wrote the implementation. This is the same discipline good teams already use for human-written code; it matters even more here because the shared mistake is easier to make.
Check that it fits your project
AI tools do not know your team's conventions unless you tell them, and even then they move away from them over time. Code can be technically correct and still be wrong for your project: it might not follow your error-handling pattern, might introduce a new dependency where you already have one that does the same job, or might structure logic in a way that conflicts with your existing architecture. Checking that code fits the project, as well as checking that it works, is what keeps a codebase coherent as more of it is written this way.
Prioritize what you review most closely
Not every line of code carries equal risk, and treating all of it with the same intensity is how review becomes exhausting and inconsistent. Read the code that touches authentication, payments, and user data with the most care, since that is where a mistake costs the most. Read internal tooling and low-stakes UI code more lightly. This is not different from how good teams already prioritize human code review; it matters more here because the volume of AI-generated code can be much higher than a team is used to reviewing.
What this looks like in practice
None of this requires slowing down to the pace of writing everything by hand. It requires making review a deliberate, separate step rather than an assumed side effect of writing the code slowly. Teams that do this well get the full speed benefit of AI-generated code without taking on its hidden mistakes. Teams that skip it end up doing this same work later, under worse conditions, which is the subject of what it takes to fix a vibe-coded app.
If you want a senior team that builds this review discipline into how they use AI tools from day one, that is core to our AI development work, and the main guide covers where this fits in the larger picture. Talk to us if you want an independent review of a specific codebase.
Common questions
What is the first step in reviewing AI-generated code?
Run the automated checks first: the test suite, static analysis, and dependency or secret scanners. This catches a meaningful share of problems for almost no effort and means your careful manual read can focus on what tools cannot catch.
Should I trust AI-generated tests to validate AI-generated code?
Not on their own. When the same AI tool generates both the implementation and its tests, the tests often validate the implementation's own understanding rather than the actual requirement. Check both against the original requirement independently, ideally with a separate person reviewing the tests.
How is reviewing AI-generated code different from reviewing a colleague's code?
The failure modes are subtler. AI output tends to look reasonable while making a wrong assumption underneath, so the review needs to focus on what the code assumes about its input, caller, and environment, not just whether it looks correct at a glance.
How long should a review of AI-generated code take?
Long enough to run the automated checks first and then read the code that touches authentication, payments, or user data closely for its assumptions. Low-stakes internal tooling can be read more lightly. There is no fixed time; the review should match what a mistake in that specific code would cost.
Do I need special tools to review AI-generated code?
The same tools that support any good code review work here: a test suite, static analysis, and dependency or secret scanning. What changes is not the tooling but the reading discipline, since AI output needs to be checked for hidden assumptions rather than skimmed for whether it looks reasonable.
Is it safe to have the same AI review its own generated code?
No, not as the only check. An AI reviewing its own output, or grading tests it also wrote, tends to repeat the same mistake rather than catch it, because both came from the same understanding of the problem. A person still needs to check the code and its tests against the original requirement.
What happens if a team skips reviewing AI-generated code?
The code tends to work in a demo and then fail on the inputs, load, or attacks a demo never tests, since those are exactly the cases a quick read would have caught. Skipping review does not remove the risk, it defers the discovery of that risk to production, where it is more expensive to fix.
Should review effort be the same for every part of an AI-generated codebase?
No. Code touching authentication, payments, and user data deserves the closest read, since a mistake there costs the most. Internal tooling and low-stakes UI code can be reviewed more lightly. Treating every line with equal intensity makes review exhausting and less consistent, not more thorough.
More in Fixing and reviewing
What it takes to fix a vibe-coded app
A vibe-coded app that mostly works does not need a rewrite. It needs its problems sorted by risk: find what is actually broken versus what is just unfamiliar, fix the parts that touch money and data first, and add the tests that were never written. Here is how that process actually goes.
Can your team maintain AI-written software?
The question that decides whether an AI-assisted build was worth it is not whether it was released. It is whether your own engineers can change it six months later without calling the people who built it. Generated code makes that question more important, because a large codebase can now be produced faster than anyone can understand it. A handover that transfers only the code transfers the least valuable part.