From AI-generated code to production software
AI can write most of a codebase in an afternoon. It cannot tell you whether that codebase is safe to run. This guide is about the difference between the two, and how to deal with it.
Published July 28, 2026. Editorial.
Key takeaways
- AI-native development means AI does most of the typing while a senior engineer owns the architecture and the risk. It is not the same as letting AI run unsupervised.
- Most AI-generated code that reaches production was never reviewed the way code from a new hire would be. The real risk is in that missing review.
- A codebase built fast with AI tools usually works in normal use and fails without warning in other cases: sign-in, unusual cases, load, and anything an attacker would try first.
- Fixing a vibe-coded app is not a rewrite. It is sorting problems by risk: find what is actually broken, secure the parts that touch money or data first, and add the tests that were never written.
- The teams that get the most value from AI coding tools are the ones with a senior engineer checking the output, not the ones with the fewest engineers.
A year ago, the question we got from founders was whether AI could write real code. That question is settled. It can, and fast. The question we get now is different: we released something built with Cursor, Lovable, or Replit, it mostly works, and we do not know what the code actually does. That question is the reason this guide exists.
Here is the short version. AI tools have made the first version of almost anything trivially fast to produce. They have not changed what it takes to run software safely for real users, with real money and real data moving through it. That work, reviewing what the AI wrote, fixing what it missed, and building the judgment to know which gaps matter, is still a job for a senior engineer. The teams that treat AI as a fast junior developer, not an autonomous one, are the ones getting real value from it. This guide is about doing that well, whether you are starting a build today or trying to understand what you already have.
What AI-native development actually means
"AI-native" gets used loosely, so it is worth being precise. It does not mean AI decides what to build or releases code without a human checking its work. It means AI is part of the architecture of how a team builds, from day one, instead of being added later as an autocomplete tool. The AI writes a large share of the code. A person still owns the intent, the architecture, and the decision about whether the output is correct. We explain the definition in more detail, and the difference between AI-native and AI-unsupervised, in what is AI-native software development.
That distinction is the whole guide in one sentence. AI-native is fast engineering that a human directs. Vibe coding without review is something else, and it is the thing that causes problems for companies.
Why the risk is real, not theoretical
We are not writing this guide because AI-generated code is bad. We are writing it because AI-generated code that nobody reviewed is now running in production at a lot of companies, and the failure pattern is consistent. It works in the demo. It works for the first real users. Then it receives an input nobody tested, a load pattern nobody planned for, or a request that is an attack, and it breaks in a way that a reviewed codebase usually would not.
The pattern shows up most in security. AI models tend to reproduce the same kinds of flaws that show up across the code they were trained on, and when that output goes straight to production without a human checking it, those flaws go with it. Secrets get left in code. Input is not validated the way it needs to be. None of this is unique to AI-written code; it is the same class of mistake a rushed human engineer makes. The difference is scale and speed: AI produces this volume of code far faster than any team could review it manually, which is exactly why review has to be deliberate rather than assumed. We describe this risk in detail in the real risks of taking vibe-coded software to production.
The review most teams skip
The single most useful habit we can recommend is treating AI-generated code the way you would treat code from a new hire you have not worked with yet: read it carefully, ask what assumptions it is making, and check that those assumptions hold for your actual system. Most teams do not do this. They read AI output the way they skim a summary, checking that it looks reasonable rather than checking that it is correct.
This is a learnable discipline, not a personality trait. It means running the tests and the static analysis first, then reading the logic with a specific question in mind: what is this code assuming about the input, the caller, and the environment, and is that assumption actually true here. We turn this into a concrete process in how to review AI-generated code before release.
Human review is only half the answer, and it is the half that cannot keep up as the volume grows. A person reads at the same speed whatever the volume, so the other half has to be automated: checks written from the specification that run on every change and can block a merge. That is a full topic of its own, and it is in the companion guide on eval-driven development.
Review is also only one step in a delivery process, and the rest of that process has to change too: how work is sized before an agent starts, which checks block a merge, how a change is released and operated, and what a leader measures. That is the AI-native delivery guide.
Fixing what is already released
A lot of teams reading this guide are not starting fresh. They already have a product that was built fast with AI tools, it mostly works, and they need to know what it will take to make it safe to keep adding to. The answer is almost never a full rewrite, and treating it as one wastes money and time you do not have.
The right approach is to sort the problems by risk. Find out what is actually broken versus what merely looks unfamiliar. Secure the parts that touch money, credentials, or user data first, since that is where a real failure costs the most. Add the tests that were never written, starting with the flows that matter most to the business. This is careful, slow work, and it is a different skill from writing new features, which is why teams that try to do it with the same fast, careless approach that produced the code often make it worse. We lay out the process in what it takes to fix a vibe-coded app.
Governance is not optional once a team grows
A single founder moving fast with AI tools is one thing. An engineering team of any real size doing the same thing without a shared standard is a different risk entirely, because now the inconsistency adds up across many people and nobody can say with confidence what is actually running. Enterprise teams adopting AI coding agents need the same basic controls they would put around any powerful new capability: who can approve a change that touches production, what gets logged, and what happens when an agent's output cannot be traced back to a person who is accountable for it. We cover what a reasonable minimum set of rules looks like in AI coding governance for enterprise teams.
A worked example of the pattern
It helps to see the failure pattern in one concrete case rather than as a general warning, so here is what it looks like on an ordinary feature.
Say a team builds a support ticket page with an AI coding tool. A customer submits a message, the message is stored, and it is later displayed to a support agent in a dashboard. The happy path works on the first try: type a message, see it appear, close the ticket. It passes the demo. It passes the first few weeks of real use, because most customers type ordinary text.
Then one customer pastes something else into the message box: a script tag, copied from a browser console after a frustrating support call, not even meant as an attack. If the message is rendered into the agent's dashboard without being escaped first, that script now runs in the agent's browser, with the agent's own session and permissions. This is cross-site scripting, and it is exactly the category Veracode's spring 2026 testing found the hardest for AI models to get right: correct 15 percent of the time, against 82 percent for a more mechanical problem like SQL injection. The reason lines up with the earlier point about context. Escaping a value correctly depends on knowing where it came from and where it is about to be displayed, and a model generating the display code often cannot see the code that stored the value, so it has no way to know the input was never trusted in the first place.
Nothing in this example required a skilled attacker. It required one customer typing an unusual string into an ordinary text box, and a rendering path that nobody had written a rule for yet. That is the shape of the risk this guide keeps coming back to: not a dramatic breach, but an ordinary input meeting a gap that the demo never exercised.
The fix is not "review harder." It is a check that runs on every change and fails the build when untrusted input reaches a rendering path without being escaped, so the gap gets caught before a real customer finds it rather than after. That is the same idea covered in is AI-generated code secure and in eval-driven development: a check written from the specification, not from the code, that can actually fail.
Why the checks have to exist before the code does
There is a specific reason a check written after the code is weaker than one written before it, and it is worth being direct about, because it explains why having tests is not the same claim as having a suite that catches real mistakes.
A check written by looking at the code that already exists tends to describe what the code does, not what it should do. If the display code forgot to escape a value, a check written by reading that same code will not notice the value was never escaped, because from the code's point of view that is simply how the feature works. The check and the mistake come from the same source, so the check agrees with the mistake.
A check written from the specification, before anyone looks at the generated implementation, asks a different question: given what this feature is supposed to do, what would have to be true for it to be safe. Untrusted input must be escaped before rendering. A support agent's session must not be usable to act on another customer's account. Those statements exist independently of whatever the AI happened to generate, so a check built from them can catch a mistake the implementation makes, rather than only confirming the implementation is internally consistent with itself.
This is the same reason a suite that has never failed is not proof of anything. A useful eval has to be shown failing on a deliberate break before anyone trusts it to catch a real one, a point covered in more detail in the eval-driven development guide. The habit worth building is writing the check, or at minimum the rule the check will enforce, at the same time the feature is specified, not after the code that was generated from that specification has already shipped.
What changes at the pace this creates
Adopting AI coding tools changes more than how fast the first version gets written, and the effect is measurable rather than a matter of impression. Google's DORA programme surveyed close to five thousand technology professionals for its 2025 report and found that AI adoption has a positive relationship with how much a team ships and a negative relationship with how stable that shipping is. More change moves through the pipeline, and a higher share of it causes problems once it is out.
Read plainly, that finding describes what has to be added alongside AI coding tools for the first half of it, the throughput gain, to be worth having without the second half, the stability cost, eating the gain. A team generating code faster than before needs its checks, its rollback path, and its incident response to be faster and more automatic than before too, or the extra throughput mostly turns into extra incidents. This is the practical reason review and governance are not optional extras layered on top of AI-native development. They are the part of the system that keeps the speed from costing more than it saves. How to move an existing engineering team to AI-native development covers what to put in place, and in what order, before turning the volume up.
The specification becomes the valuable part
One more shift is worth naming on its own, because it changes what is actually worth protecting in a build.
On a codebase written by hand, the code itself is the expensive, hard-won artefact, and everything else, the design notes, the comments, the tickets, describes it after the fact. Once AI is doing a large share of the typing, that relationship flips. The implementation becomes the cheap, regenerable part. The precise statement of what the software must do, the specification a check can be built from and the implementation can be checked against, becomes the part that took the real thinking and the part a team cannot afford to lose.
This shows up directly at handover. A team that hands over a repository and calls the job finished has given you the part that was fast to produce and kept the part that explains why it works the way it does. A team that hands over the specification and the eval suite alongside the code has given you something a new engineer can actually check their own understanding against, rather than something they have to reverse-engineer line by line. Can your team maintain AI-written software sets out what a handover needs to contain and how to test it before accepting the work, rather than discovering the gap the first time your own team tries to change something.
Choosing who helps you
Whether you need to review a codebase that already exists, or you want a partner who builds AI-native from day one and does the review as part of the work, the team you hire matters more than the tools they use. A partner who tells you honestly what state your code is in, rather than one who is eager to bill hours rewriting it, is worth finding before you commit. We cover what to ask in how to evaluate an AI development partner.
Where Reveneau fits
This is the work Reveneau is built around. A prototype built in Lovable or Cursor proved the idea and now has to become something real, or a team moved fast with AI tools and needs the output reviewed and secured before real users depend on it. Either way the job is the same: keep the speed AI gives you, and add the judgment that speed alone does not include. That is what our AI development work is built around, and it is also most of what a full product build looks like when the starting point is a fast AI-built prototype rather than an empty codebase.
If your team is already deciding whether AI is a feature or the whole product, that is a different and earlier question, covered in the building AI products guide. If the question you face right now is whether what you already have is safe to keep adding to, talk to us and we will tell you honestly.
Explore the guide
Understanding the risk
What is AI-native software development?
AI-native software development means AI is part of how a team builds from the start, writing a large share of the code, while a person still owns the architecture, the intent, and the decision about whether the output is correct. It is not the same as AI building without anyone checking.
The real risks of taking vibe-coded software to production
Vibe-coded software that never gets a real review tends to fail in a specific, predictable way: it works in normal use and breaks in every other case, including the places an attacker would look first. Here is what that risk actually looks like and why it is not theoretical.
Is AI-generated code secure? What the testing shows
Independent testing of more than 150 language models found that only 55 percent of their code generations were secure when no security guidance was given. That number has barely changed since 2023, while the models' ability to produce code that simply runs has risen to around 95 percent. The gap between code that works and code that is safe is getting wider, and that is the single most important fact to keep in mind when you decide how AI-written code reaches your customers.
Fixing and reviewing
How to review AI-generated code before release
AI-generated code should get a closer read than code from a colleague you trust, not a lighter one, because the failure modes are subtler and speed makes it tempting to skim. Here is a concrete process for reviewing it properly.
What it takes to fix a vibe-coded app
A vibe-coded app that mostly works does not need a rewrite. It needs its problems sorted by risk: find what is actually broken versus what is just unfamiliar, fix the parts that touch money and data first, and add the tests that were never written. Here is how that process actually goes.
Can your team maintain AI-written software?
The question that decides whether an AI-assisted build was worth it is not whether it was released. It is whether your own engineers can change it six months later without calling the people who built it. Generated code makes that question more important, because a large codebase can now be produced faster than anyone can understand it. A handover that transfers only the code transfers the least valuable part.
Choosing who helps you
AI coding governance: what enterprise teams need in place
One engineer moving fast with AI tools is a personal workflow choice. A whole team doing it without a shared standard is a governance gap, because the inconsistency adds up across every person changing the code. Here is what a reasonable minimum set of controls looks like.
Who reviews AI-generated code, if there is no human sign-off?
An automated check can only catch a mistake somebody already wrote a rule for. It cannot notice that the rule itself is missing. That is the real question behind "who reviews AI-generated code": not whether a check ran, but who is responsible when the failure is a rule nobody thought of yet.
How to evaluate an AI development partner
Hiring an AI development partner is harder than a normal build decision, because you are not buying a finished capability, you are buying judgment applied to something that does not fully exist yet. Here is what actually separates a good partner from a confident pitch.
Common questions
What does AI-native software development mean?
It means AI writes a large share of the code as part of how a team builds from day one, while a person still owns the architecture, the intent, and the decision about whether the output is correct. It is not the same as AI releasing code without a human checking its work.
Is AI-generated code safe to put in production?
Only after it has been reviewed the way you would review code from a new engineer you have not worked with yet. Unreviewed AI output tends to work in normal use and fail on unusual cases, load, and any input that is an attack, because those are exactly the cases a demo never tests.
How do I fix an app that was built with AI coding tools?
Sort the problems by risk rather than rewrite. Find out what is actually broken versus what is just unfamiliar, secure the parts that touch money, credentials, or user data first, and add the tests that were never written. A full rewrite is rarely necessary and usually wastes the speed you have already gained.
Do AI coding agents need governance in an enterprise setting?
Yes, once more than one person is using them. The basics are the same as any powerful capability: who can approve a change that touches production, what gets logged, and whether an agent's output can be traced back to a person who is accountable for it.
Should I hire a team that uses AI coding tools?
The tools matter less than the review discipline around them. Ask any potential partner how they review AI-generated output before it is released, not just whether they use AI tools. A partner who is honest about the state of a codebase, rather than eager to bill hours rewriting it, is the one worth hiring.
What is the difference between AI-native development and vibe coding?
AI-native development means AI writes most of the code while a person still owns the architecture and checks the output before it is released. Vibe coding, in the risky sense this guide warns about, means accepting AI output and releasing it without that review. Both use the same tools; only one is safe to run a real business on.
How long does it take to fix a codebase built fast with AI tools?
The timeline is set by how much of the codebase touches money, credentials, or user data, since that is what gets reviewed and secured first. The process is sorting by risk, not a fixed timeline: an honest audit finds what is actually broken, then the highest-risk flows get fixed before anything cosmetic, with tests added along the way.
Is it risky to keep building on a vibe-coded app instead of fixing it first?
Yes, if the parts handling authentication, payments, or user data have never been reviewed. Adding features on top of unreviewed code increases the risk instead of resolving it. Fixing in risk order first, then continuing to build, protects the parts that would cost the most if they failed.
Related reading
An AI demo is not a product
A convincing AI demo takes an afternoon. Turning it into something people trust in production takes most of the work, and most failures happen at that stage.
Can you trust an AI agent with real work yet?
An agent that answers a question and an agent that takes an action are not the same risk. Here is how we decide where an agent is ready to act, and where it is not.
How to get a working AI prototype in weeks, not quarters
Most AI ideas end during the planning stage. Here is how to show something real to users fast enough to know if the idea is worth the full build.