From AI-generated code to production software / Understanding the risk
Is AI-generated code secure? What the testing shows
Independent testing of more than 150 language models found that only 55 percent of their code generations were secure when no security guidance was given. That number has barely changed since 2023, while the models' ability to produce code that simply runs has risen to around 95 percent. The gap between code that works and code that is safe is getting wider, and that is the single most important fact to keep in mind when you decide how AI-written code reaches your customers.
Published July 28, 2026. Editorial.
Key takeaways
- Only about 55 percent of AI code generations were secure in Veracode's spring 2026 testing, across 150-plus models and 80 tasks.
- The security figure has stayed between 45 and 55 percent since 2023 while syntax correctness rose from roughly 50 to 95 percent.
- Results vary sharply by language and flaw type: Java passed 29 percent of the time, cross-site scripting 15 percent, log injection 13 percent.
- Some categories are handled well, including SQL injection at 82 percent and cryptographic algorithm choice at 86 percent.
- Reading every diff does not scale as a security control, so the check has to be automated and has to block a merge.
Ask whether AI-generated code is secure and you usually get an opinion. There is measured evidence now, and it is specific enough to plan around.
The short answer
Veracode's spring 2026 study tested more than 150 large language models against 80 coding tasks, across four languages and four vulnerability types. Only 55 percent of the generations produced secure code when the prompt gave no explicit security guidance [1]. Put plainly, close to half of what came back contained a known vulnerability.
The more useful number is the trend. Security performance has stayed between 45 and 55 percent since 2023. Over the same period, syntax correctness rose from roughly 50 percent to about 95 percent. Models got better at writing code that compiles and runs, and did not get better at writing code that is safe. Veracode describes the gap between code that works and code that works securely as growing wider.
That matters because the thing that improved is the thing you notice, and the thing that did not improve is the thing you do not.
Where it fails, and where it does not
The averages hide large differences. Broken out by language, Python passed 62 percent of the time, C# 58 percent, JavaScript 57 percent, and Java 29 percent. A Java service built this way is in a different risk category from a Python one, and no general statement about "AI code quality" stays true across results that far apart.
By vulnerability type the range is wider. SQL injection was handled correctly 82 percent of the time and insecure cryptographic algorithms 86 percent. Cross-site scripting was handled correctly in 15 percent of cases, and log injection in 13 percent.
There is a pattern worth naming. The categories the models handle well are the ones with a single standard fix that appears constantly in public code: use a parameterised query, use a modern cipher. The categories they handle badly are the ones where safety depends on context the model cannot see, such as where a value came from, whether it has already been escaped, and where it is about to be rendered. Cross-site scripting is exactly that kind of problem, and it is the one the models are worst at.
So the risk is not spread evenly across your codebase. It is concentrated in the code that takes untrusted input and passes it to a place where it will be interpreted.
Why "we review everything" is a weak answer
The obvious control is human review, and it is much weaker here than it feels.
The first reason is volume. Generated code arrives faster than anyone reads it, and review quality falls as diff size grows. A reviewer facing 900 lines does not apply the attention they would give 90.
The second reason is that these flaws are, by accident, the kind a reader misses. Generated code is fluent. It follows conventions, it is commented, it looks like the surrounding code. A missing escape on one path in a well-structured file is harder to see than a badly written function.
The third reason is that we are poor judges of our own performance with these tools. In a randomised controlled trial published in July 2025, METR gave 16 experienced open-source developers 246 real tasks in repositories they already knew well. The developers using AI tools took 19 percent longer, and afterwards estimated that AI had made them 20 percent faster [2]. METR is careful to label the result as historical, since the tools have changed since then, and that caveat is fair. The finding that still holds is that skilled people misjudged their own performance by 39 points, in the direction that made the tool look better. If your sense of your own speed is that unreliable, your sense of whether a diff is safe is not going to be better.
Developers seem to sense this. In Stack Overflow's 2025 survey, 46 percent said they actively distrust the accuracy of AI output, against 33 percent who trust it [3].
What actually works
Security has to move from something a person notices to something a pipeline enforces. Three things matter most.
Static analysis that blocks the merge. Not a report someone reads on Friday. A check that fails the build. If it only warns, it will be ignored within a month.
Checks aimed at the categories that fail. The general scanner is the minimum. Add explicit assertions on the paths where untrusted input reaches a place where it is used: rendering to a page, writing to a log, building a query, constructing a shell command. Those are where the measured failure rate is worst, and a targeted check finds what a general one misses.
Prompt-level guidance, with verification anyway. The 55 percent figure is for generations with no security guidance in the prompt. Asking for secure output helps, and it is free, so do it. It is not a control, because you cannot tell from the output whether it worked. Ask, then verify.
All three follow the same idea covered across this guide: the check has to exist before the code, and it has to be able to fail. That is the subject of eval-driven development, and the security case is the clearest argument for it.
What to ask a vendor
If someone is writing AI-generated code for you, four questions separate a real answer from a reassuring one.
Which security checks run on every change, and do they block a merge or produce a report? What happens to a pull request that fails one? Which languages and flaw classes are you treating as higher risk, and why those? When a vulnerability reaches production, what changes in the pipeline afterwards, not just in the code?
The last one matters most. A team that fixes the bug and does nothing more will see the same bug again. A team that adds the check that would have caught it gets better protection with each bug. We wrote about the cost of skipping that step in the real cost of shipping unverified code.
If you want the wider view, how to review AI-generated code covers what a person should still look at, and AI coding governance covers the rules once more than one team is building this way. If you are deciding whether to generate the whole thing yourself instead, Reveneau vs an AI app builder sets out where that limit is.
Best for
- Teams releasing AI-generated code who want the security question settled with evidence rather than opinion
- Anyone choosing where to spend limited security effort across a large generated codebase
- Buyers writing security requirements into a contract with a development partner
Avoid if
- Do not read the 55 percent figure as a reason to stop using AI to write code: the same testing shows syntax correctness near 95 percent
- Do not treat a passing general-purpose scanner as coverage for the categories that fail worst
- Do not rely on asking the model for secure code, since you cannot tell from the output whether the request worked
Check before you decide
- Confirm a security check blocks the merge rather than filing a report someone reads later
- Confirm the paths where untrusted input reaches a renderer, a log, or a query carry explicit assertions
- Confirm each vulnerability that reaches production produces a new automated check, not only a patch
Common questions
What percentage of AI-generated code is secure?
In Veracode's spring 2026 testing of more than 150 models across 80 coding tasks, 55 percent of generations produced secure code when the prompt included no security guidance. The figure has stayed between 45 and 55 percent since 2023 while syntax correctness rose to around 95 percent.
Which languages produce the least secure AI-generated code?
Java was the clear outlier in Veracode's testing at a 29 percent security pass rate, against 62 percent for Python, 58 percent for C#, and 57 percent for JavaScript. A Java service generated this way needs materially more verification than a Python one.
Which vulnerability types do models handle worst?
Log injection at 13 percent and cross-site scripting at 15 percent were the weakest categories. Both depend on context the model cannot see, such as where a value came from and whether it was already escaped, which is why a fluent-looking answer often gets them wrong.
Are there security categories AI handles well?
Yes, and it is worth knowing which. SQL injection was handled correctly 82 percent of the time and insecure cryptographic algorithm choice 86 percent. Those are problems with one standard fix that appears constantly in public code, which is the kind of pattern models learn reliably.
Can code review catch these vulnerabilities?
Not dependably at volume. Generated code is fluent and conventional, so a missing escape on one path does not stand out the way sloppy code would, and review quality drops as diff size grows. Review is worth doing for design and intent, but it is the wrong primary control for this class of flaw.
Does asking the model for secure code fix the problem?
It helps and it costs nothing, so include it. It is not a control, because the output gives you no signal about whether the instruction was followed. The 55 percent figure applies to generations without guidance; treat guidance as an improvement to the input, then verify the output regardless.
What should a security check on AI-generated code actually do?
It should fail the build rather than produce a report. Add targeted assertions on the paths where untrusted input reaches a page, a log, a query, or a shell command, since those map to the categories with the worst measured pass rates. A warning that does not block gets ignored within about a month.
How do I make a development partner meet this standard?
Ask which checks block a merge, what happens to a pull request that fails one, and what changes in the pipeline after a vulnerability reaches production. The last question is the most revealing one: fixing the bug alone means meeting it again, while adding the check that would have caught it means the suite gets stronger with each incident.
References
- Veracode, Spring 2026 GenAI Code Security Report: 55% of generations secure across 150+ models and 80 tasks; Java 29%, Python 62%, C# 58%, JavaScript 57%; XSS 15%, log injection 13%, SQL injection 82%, insecure crypto 86%; security flat at 45 to 55% since 2023 while syntax correctness rose from ~50% to ~95%.
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (July 2025): 16 developers, 246 tasks, 19% longer completion times against a self-estimated 20% speed-up. METR labels the result historical.
- Stack Overflow 2025 Developer Survey, AI section: 46% actively distrust the accuracy of AI output against 33% who trust it.
Related reading
The real cost of shipping unverified code
The cost of unverified code does not arrive as a bug report. It arrives as a codebase nobody will touch, a review queue that never empties, and a team that has stopped trusting its own pipeline.
A practical pre-launch security review for a small team
You do not need perfect security to launch. You need to check the few basics that find most real problems, and to know when the risk is big enough to bring in a specialist.
How many tests does AI-generated code need?
The honest answer is not a number, and coverage percentages are the wrong unit. Here is the unit we use instead, and why a suite of fifteen checks can be worth more than two thousand.
More in Understanding the risk
What is AI-native software development?
AI-native software development means AI is part of how a team builds from the start, writing a large share of the code, while a person still owns the architecture, the intent, and the decision about whether the output is correct. It is not the same as AI building without anyone checking.
The real risks of taking vibe-coded software to production
Vibe-coded software that never gets a real review tends to fail in a specific, predictable way: it works in normal use and breaks in every other case, including the places an attacker would look first. Here is what that risk actually looks like and why it is not theoretical.