From AI-generated code to production software / Fixing and reviewing
Can your team maintain AI-written software?
The question that decides whether an AI-assisted build was worth it is not whether it was released. It is whether your own engineers can change it six months later without calling the people who built it. Generated code makes that question more important, because a large codebase can now be produced faster than anyone can understand it. A handover that transfers only the code transfers the least valuable part.
Published July 28, 2026. Editorial.
Key takeaways
- If your team cannot safely change the software without the original builder, you bought a dependency rather than an asset.
- The code is the least valuable thing in a handover. The specification, the eval suite, and the decision record contain the understanding.
- Generated code tends to become inconsistent from one part to another, because each generation sees only the context it was given.
- Test a handover before you accept it: have your own engineer make a real change while the original team watches without helping.
- Put the artefacts in the contract, since asking for them after the final invoice rarely works.
Software gets handed over twice. Once on paper, when the invoice is settled, and once in reality, the first time someone on your side has to change it under pressure. The second handover is the one that counts, and plenty of projects that passed the first fail the second.
AI-written code makes both handovers more important. A working system can be produced far faster than before, which is the point. But the understanding of why it is shaped the way it is does not arrive at the same speed, and it is not in the files.
The test that matters
Here is the only test worth running. Take a real change, not a small practice one. Something a user would notice, touching more than one part of the system. Give it to an engineer on your side who did not build the thing. Let them do it while the original team watches and answers nothing.
If they can do it, you own the software. If they cannot, you depend on whoever built it for every change, whatever the contract says about intellectual property.
This is uncomfortable to run, which is exactly why it is informative. Everyone agrees the handover went well until someone tries to use it.
Why generated code drifts
Handover difficulty is not new. Two things about generated code make it worse, and both are structural rather than a matter of care.
The first is local consistency. A model works from the context it is given. Across a large codebase built over weeks, different parts get generated with different context, so the same problem gets solved three reasonable ways in three places. None of them is wrong. Together they mean a new reader cannot form one understanding and apply it everywhere, which is the thing that makes a codebase fast to work in.
The second is missing intent. Generated code is unusually good at explaining what it does, because comments and naming are exactly what models are good at. It says nothing about what was rejected. Why this approach and not the obvious alternative, which constraint forced the awkward design, what broke last time. That information never existed in text, so it cannot be in the repository. It was in a conversation.
The volume of code adds to both problems. Google's DORA programme surveyed nearly 5,000 technology professionals for its 2025 report and found AI adoption has a positive relationship with delivery throughput and a negative relationship with delivery stability [1]. More change arriving, a higher share of it going wrong. A team inheriting that needs more support than a team inheriting a smaller, slower codebase, not less.
What a real handover contains
Four things, in rough order of how much they matter.
The specification the software was built from. On an AI-native build this is not documentation written afterwards. It is the input the implementation was generated against, so it is the closest thing to a statement of intent that actually shaped the code. If it has been kept current, it is worth more than the code.
The eval suite, and proof it can fail. A new maintainer's first question is "if I change this, will I know I broke something." The suite is the answer. It has to be runnable on their machine on day one, and someone should show that a deliberate break makes the build fail. A suite that has always passed and never failed proves nothing. That is covered in when evals give false confidence.
A short record of the decisions. Not a design document. A list of the choices that would look wrong to a newcomer, each with the reason and the alternative that was rejected. Ten entries of three sentences beats forty pages nobody opens.
A runbook for the routine emergencies. How it deploys, how to roll back, where the logs are, what the alerts mean, which failures are expected and safe to ignore. This is the artefact people skip and then need at two in the morning.
The specification is the asset now
There is a reversal here worth stating on its own, because it changes what you should be negotiating over.
On a hand-written project, the code is the artefact. It holds the effort, and everything else describes it after the fact. On an AI-native project the relationship reverses. The implementation is the cheap, regenerable part. The expensive, difficult part is the precise statement of what the software must do, because that is what took the arguing, and it is what any future change has to be checked against.
That has a practical consequence. If you have the specification and the eval suite, a large part of the implementation can be regenerated. If you have only the implementation, you have the one thing that was easy to produce and none of the reasoning that made it correct. A team that hands over a repository and calls it done has given you the easy part and kept the valuable part.
It also explains why the specification has to be current rather than the version written before the work started. Requirements change during a build. A specification that stopped being true in week two is a historical document, not a maintainable asset, and the gap between it and the code is exactly where the next maintainer will make mistakes.
Getting it, rather than hoping for it
Ask for these while you can still insist on them, which means in the contract, not after the last invoice.
Three clauses matter most. Name the artefacts as deliverables in their own right, so the work is not complete without them. Make the handover test above an acceptance condition, with a named engineer on your side and a real change. And schedule a support window after the handover that is deliberately consultative, where the original team answers questions rather than making changes, because the moment they start making the changes again you depend on them again.
One thing worth saying plainly: a team that has built this way properly will welcome all three, because the artefacts already exist as a by-product of how they work. A team that resists is usually telling you the specification was never written down and the suite is weak. That reaction is itself a useful answer.
The honest limit
None of this makes an unfamiliar codebase feel familiar. Your engineers will still be slower in it for a while, and expecting otherwise will annoy everyone. What a good handover buys is that the slowness is temporary and self-resolving, rather than a permanent cost that sends you back to the original team for every change.
If you are deciding what to keep in-house in the first place, in-house vs outsourced engineering covers that split, and how to review AI-generated code covers what your reviewers should be looking at while the work is still in progress.
Best for
- Buyers taking delivery of a build where AI wrote most of the implementation
- Teams inheriting a codebase they did not write and have to own from now on
- Anyone drafting the acceptance terms for an AI-assisted engagement
Avoid if
- Do not run the handover test as a formality with the original team helping: a supervised pass proves nothing
- Do not accept documentation written after the fact in place of the specification the code was generated from
- Do not expect your engineers to work at full speed immediately in a codebase they did not build
Check before you decide
- Confirm an engineer on your side can make and release a real, user-visible change unaided
- Confirm the eval suite runs on a fresh machine and fails when you break something on purpose
- Confirm the artefacts are named as contractual deliverables, not promised verbally
Common questions
How do I know if my team can actually maintain AI-written software?
Give an engineer who did not build it a real, user-visible change and have the original team watch without helping. If they complete the change, you own the software. If it does not, you own a dependency on the builder regardless of what the contract says about ownership.
What should a handover include besides the code?
The specification the code was generated from, the eval suite with proof it can fail, a short record of the decisions that would look wrong to a newcomer, and a runbook for deploys, rollbacks, logs, and alerts. The code is the least valuable item on that list because it is the one thing you can always read.
Why is AI-generated code harder to inherit than hand-written code?
Two structural reasons. Each generation sees only the context it was given, so the same problem gets solved several reasonable ways across a large codebase, and no single understanding covers all of it. And generated code explains what it does while saying nothing about what was rejected and why, because that reasoning existed only in a conversation rather than in the repository.
Is a passing test suite enough proof that a handover is safe?
No, because a suite that has never failed may be asserting nothing. Ask someone to break a protected behaviour on purpose in front of you and confirm the build fails. Checks that mock away the real path or compare a value to itself are common enough that this demonstration is worth the five minutes.
When should I ask for handover artefacts?
In the contract, before work starts. Name them as deliverables so the engagement is not complete without them, and make the handover test an acceptance condition. If you ask after the final invoice, you no longer have a reliable way to get them.
What does it mean if a partner resists these terms?
Usually that the specification was never written down and the eval suite is weaker than described. A team that builds this way properly already produces these artefacts as a by-product, so the request costs them little. The reaction is a cheap and honest signal about how the work is actually being done.
How long should the support window after handover be?
Long enough to cover a full release cycle and at least one production incident, and structured so the original team answers questions rather than making the changes. The moment they start making changes again, the transfer of understanding stops and you depend on them again.
Does AI-written software need a different handover than hand-written software?
Yes, because the risk changes. A hand-written codebase mainly needs documentation of intent. AI-written code adds a second problem: different parts can be generated with different context, so the same problem gets solved several reasonable ways across the codebase, which makes a single understanding harder to form without the original specification.
Related reading
In-house vs outsourced engineering: what to keep and what to give to others
The real question is not whether to build in-house or outsource. It is which parts belong to your own team forever, and which parts an outside team can do faster and better right now.
What good code review looks like when nobody wrote the code
With human code, the author is the first check and review is the second. With generated code, review is the only check. That one change alters most of what a reviewer should be doing.
How to onboard engineers fast so they release work in week one
Slow onboarding is a hidden cost you pay on every hire. Here is how to get a new engineer releasing real work in their first week instead of their first month.
More in Fixing and reviewing
How to review AI-generated code before release
AI-generated code should get a closer read than code from a colleague you trust, not a lighter one, because the failure modes are subtler and speed makes it tempting to skim. Here is a concrete process for reviewing it properly.
What it takes to fix a vibe-coded app
A vibe-coded app that mostly works does not need a rewrite. It needs its problems sorted by risk: find what is actually broken versus what is just unfamiliar, fix the parts that touch money and data first, and add the tests that were never written. Here is how that process actually goes.