Technical due diligence: a complete guide for investors / What to assess
How to assess code quality
Code quality in technical due diligence means one thing: whether the software is safe and cheap to change. If small changes are slow and risky, the roadmap in the deal deck will not be delivered on time. The most reliable evidence comes from how the team releases software today, measured with the five DORA delivery metrics, a live run of the test suite, and the pull request history, rather than from how the code looks in a screenshot. This page explains what to measure, what to ask, and how to judge quality when you cannot read the code yourself.
Published July 27, 2026. Updated September 17, 2026. Editorial.
Key takeaways
- Code quality means the software is safe and cheap to change, and the strongest signal is how fast and safely the team releases a small change to production today.
- DORA's five delivery metrics (change lead time, deployment frequency, failed deployment recovery time, change fail rate, deployment rework rate) turn quality into numbers a deal team can compare.
- Automated tests that run in front of the reviewer, a working release process, and small reviewed changes matter more than any style metric.
- AI-written code is normal in 2026 and is a finding only when it arrives without a review record and a verification method the team can show running.
When investors hear "code quality" they often picture clean, elegant code. Quality in the context of a deal means something more practical: can the team change this software quickly and safely, or does every change risk breaking something. If the answer is the second one, the roadmap being funded will be delivered far slower than the plan promises.
What does code quality mean in a deal?
Code quality in a deal means the software can be changed at the speed the thesis assumes, with a low chance of each change breaking production. That definition is about behaviour rather than appearance, and it can be measured from the outside. Software that is easy to change lets the team release features, fix bugs, and respond to the market. Software that is hard to change turns every small request into a slow, risky project.
So the question to answer in diligence is: how fast and how safely does this team release a small change today. A team that can put a fix into production in a day, with confidence it will not break something else, has good quality by the only definition that matters to the deal. A team where a one-line change takes a week has a problem, no matter how tidy the code looks. This connects directly to assessing technical debt, because debt is the main thing that makes change slow and risky, and to the main guide, which presents every finding as a cost against the thesis.
Reveneau's code quality assessment asks for the target's delivery numbers for the last twelve months before it asks for the code, because a team that can produce those numbers has already told the reviewer most of what a style review would, and a team that cannot produce them has told the reviewer something too.
Which numbers tell you whether code is safe to change?
The DORA delivery metrics tell you. DORA, the research programme that has published an annual study of software delivery from 2014 to 2025, defines five measures, three of throughput and two of instability. They are the closest thing the industry has to a standard set of measures, and a target that tracks them can provide them in an afternoon.
| DORA metric | DORA's definition | What it tells the deal team |
|---|---|---|
| Change lead time | The time for a change to go from committed to version control to deployed in production | How long a roadmap item waits between "done" and "live" |
| Deployment frequency | The number of deployments over a period, or the time between deployments | Whether the release process is routine or an event |
| Failed deployment recovery time | The time to recover from a deployment that fails and needs immediate intervention | How long customers wait when something breaks |
| Change fail rate | The ratio of deployments that require immediate intervention | How often a change breaks production |
| Deployment rework rate | The ratio of deployments that are unplanned and happen as a result of an incident | How much release work is emergency repair rather than planned roadmap work |
DORA's research overview also records its own main findings: high-performing organisations do better on all of the delivery measures at once, so speed and stability rise together rather than one improving at the cost of the other, and its 2023 study found teams that prioritise user needs achieve 40 percent higher organisational performance. Ask for the five numbers, ask how they were measured, and ask for the trend over twelve months. A team whose lead time is growing and whose deploy frequency is falling is a team whose code is getting harder to change.
What should you look at in the repository and the release process?
Look at four things: whether automated tests exist and run, how changes are reviewed, how large changes are, and what a release involves. Each can be checked in a working session with an engineer rather than by reading the code line by line.
Tests. Ask what share of the code is covered by automated tests, then ask to watch the suite run. A suite that passes in front of you is evidence. A suite that "usually passes" or "needs a few environment fixes" is a finding. A codebase with no tests is a codebase where every change is a risk, and the how many tests does AI-generated code need post covers what coverage means for a modern codebase.
Review. Google's published engineering practices state that reviewers should approve a change once it "definitely improves the overall code health of the system", even if it is imperfect, and that a change should not wait days or weeks for perfection. Look at the pull request history for that pattern: named reviewers, comments that show reading, approvals that take longer than the change could have been read. Approvals arriving seconds after a large change was opened tell you the review is only a formality.
Change size. Google's guidance on small changes states that 100 lines is usually a reasonable size and 1,000 lines is usually too large, because small changes are reviewed more quickly, reviewed more thoroughly, and less likely to introduce bugs. A repository where most changes are thousands of lines is a repository where review cannot have caught much.
Release. Ask what happens between "merged" and "live". A scripted pipeline that deploys many times a week is a good sign. A release that needs a named person, a weekend, and a rollback plan drawn on a whiteboard is a cost to price.
How do you judge quality without reading code?
Judge it from behaviour, because quality is visible in how the company operates. Most investors are not engineers, and the review does not need them to be.
- Ask for a walkthrough of one recent feature. A strong team tells you the options, the choice, and the trade-off in plain language. A team that answers with jargon either does not understand its own code or is hiding careless choices.
- Ask what a one-line change costs. How long from idea to production, and who has to be involved. The honest answer is a lead time, and it is the same number DORA measures.
- Ask about the last time a release broke something. How was it found, how long did recovery take, what changed afterwards. This is the change fail rate and recovery time, described in the team's own words.
- Talk to someone who inherited the code. The most revealing reference question is how the work performed after it was released: did small changes stay cheap, or did they become slow and hard. Software that is fast to build and expensive to change is a common form of poor quality that is hard to see.
The assessing the engineering team page covers assessing the people in more detail, and the same logic applies to any external team you evaluate, as in how to choose the right external development team.
What changes when the code was written by AI?
The standard changes from "who wrote this" to "what checked this". Most codebases reviewed in 2026 contain code that a model wrote, and penalising a company for it would rule out most of the teams building well. The finding is generated code with no review record and no verification method, because that means the company cannot fully explain its own system.
Three figures describe the risk. Stack Overflow's 2025 survey of more than 49,000 developers found 84 percent using or planning to use AI tools, 46 percent distrusting the accuracy of AI output against 33 percent trusting it, and 66 percent naming solutions that are "almost right" as their biggest frustration. Google's October 2024 DORA report associated a 25 percent increase in AI adoption with a 1.5 percent decrease in delivery throughput and a 7.2 percent decrease in delivery stability. Veracode's July 2025 study of more than 100 language models found 45 percent of generated code samples failed security tests.
So the question for the review is what catches the almost-right code before production. Look for specifications that are still kept up to date, for a test or evaluation suite that runs in the pipeline and fails when it should, and for a pull request history that shows human reading. Our AI-generated code to production hub covers the full argument and the security figures, the eval-driven development hub explains the evaluation suite a reviewer should ask to see, and the diligence on an AI-written codebase page covers the review itself.
How does quality connect to the team?
Most of what people call code quality depends on who wrote it and who reviewed it. Experienced engineers have seen which quick solutions cause costs later, so they avoid them without being told. Across a review, weak code and a junior team tend to appear together, and strong code and a senior team do too. When you find good quality, check that the seniority behind it is still at the company. When you find weak quality, the fix is usually as much about the team as the code. Stripe's 2018 developer survey, by its own account, found the average developer spending 17.3 hours of a 41.1-hour week on maintenance work such as debugging and refactoring, and 13.5 hours on technical debt; a team that reports numbers in that range is describing a codebase that is hard to change.
How do you put the finding in the deal?
Price it. Poor quality is a cost rather than an automatic reason to leave the deal. Estimate what it takes to get the code to where the roadmap needs it, and raise that number in the deal negotiation.
| Finding | Typical repair | Where it goes |
|---|---|---|
| No automated tests, but a senior team | Build a suite around the main functions first | First-hundred-days plan |
| Long lead time, rare releases | Pipeline and release automation | First-hundred-days plan |
| Large unreviewed changes, approvals that are only a formality | Review policy, smaller changes, a reviewer who reads | First-hundred-days plan, and a team finding |
| AI-written code with no verification method | Specification recovery, an evaluation suite in the pipeline | Price, because the risk is unmeasured |
| Team cannot change its own software safely and does not know it | Senior hires or a partner team | Terms, and a warning sign |
The technical due diligence checklist lists the code quality questions in a form you can take into a review, and the report structure and template page shows how a priced finding is written up.
Best for
- Judging whether the roadmap in the deck can be delivered at the speed the thesis assumes
- Reviews where the investor cannot read code and needs behavioural evidence instead
- Any codebase where a large share of the code was written by AI
Avoid if
- You need a security review, which is a separate page and a separate discipline
- The deal thesis does not depend on release speed at all, which is rare for software
Check before you decide
- Confirm the five DORA numbers for the last twelve months and how they were measured
- Confirm the test suite passes in front of the reviewer, on the reviewer's schedule
- Confirm the pull request history shows human reading rather than instant approvals
Common questions
How do you assess code quality without being an engineer?
Judge behaviour and outcomes rather than the code itself. Ask how fast and safely the team releases a small change, ask for a plain-language explanation of one recent feature, and ask someone who maintained the code whether changes stayed cheap. DORA's five delivery metrics, read on 17 September 2026, turn those questions into numbers: change lead time, deployment frequency, failed deployment recovery time, change fail rate and deployment rework rate.
What is the best single signal of code quality in due diligence?
The best single signal is how fast and safely the team releases a small change to production today. A team that can release a fix in a day with confidence has good quality by the definition that matters. DORA's research overview records that high performers do better on delivery speed and stability at the same time, so a short lead time with a low change fail rate is the pattern to look for.
Does poor code quality end a deal?
Rarely on its own. Poor code quality is a cost you price into the deal: a test suite, release automation, or a stronger team. It becomes a deal risk only when the team cannot change its own software safely and does not realise it. Stripe's 2018 survey, by its own account, found developers spending 13.5 hours a week on technical debt, so a team reporting numbers in that range is describing a repair rather than a disaster.
What is the difference between code quality and code cleanliness?
Code quality means the software is safe and cheap to change; cleanliness means it looks tidy. A codebase can look elegant and still be risky to modify, and a messy-looking codebase can be released fast and safely when it has good tests and a solid release process. Google's published code review standard asks reviewers to approve a change once it improves overall code health rather than waiting for perfection, which is the same priority.
How do automated tests relate to code quality?
Automated tests let a team change software without fear by catching mistakes before users do, which makes them one of the strongest signs of quality. Ask to watch the suite run rather than accepting a coverage number. Stack Overflow's 2025 survey found 66 percent of developers frustrated by AI output that is almost right, and a test suite that runs in the pipeline is what catches almost-right code in a modern codebase.
What questions reveal code quality in technical due diligence?
Ask for the five DORA delivery numbers over twelve months, ask to watch the tests run, ask how often the company releases and what happens when a release fails, and ask a senior engineer to explain step by step how a recent feature was built. Google's small-changes guidance, which puts a reasonable change at 100 lines and 1,000 lines as usually too large, gives you a benchmark for the pull request history too.
How does release frequency connect to code quality?
Frequent, quick releases show a working pipeline and code that supports change, while rare, difficult releases show the opposite. DORA defines deployment frequency as the number of deployments over a period or the time between them, and its research overview records that high-performing organisations improve speed and stability together rather than trading one for the other.
Is AI-written code a warning sign in technical due diligence?
No. AI-written code with no review record and no verification method is the finding, because it means the company cannot fully explain its own system. Google's October 2024 DORA report associated a 25 percent increase in AI adoption with a 7.2 percent decrease in delivery stability, so the review asks what catches the almost-right code: specifications, a test or evaluation suite in the pipeline, and a pull request history that shows reading.
What does a pull request history tell you about code quality?
It tells you whether review is real. Named reviewers, comments that show reading, and approvals that took longer than the change could have been read are evidence of a team that catches its own mistakes. Google's engineering practices favour changes near 100 lines because they are reviewed more quickly and thoroughly and are less likely to introduce bugs, so a history of multi-thousand-line changes approved in seconds is a finding.
Should you verify code quality claims or trust the team's word?
Verify them. Ask to see the tests run and talk to someone who maintained the code afterwards, because software that is fast to build and expensive to change is a common form of poor quality that is hard to see. Veracode's July 2025 study found 45 percent of AI-generated code samples failed security tests, so a claim that the code is well tested needs a test suite running in front of the reviewer rather than a number on a slide.
How do you price a code quality finding in a deal?
Estimate the work to get the code to where the roadmap needs it and raise that number in the price negotiation. A missing test suite is a fixed piece of work for the first hundred days. Long lead times are a pipeline project. AI-written code with no verification method is priced as unmeasured risk. Stripe's 2018 survey figure of 17.3 hours a week on maintenance, by its own account, gives a sense of what a repair that is never made costs in team time.
References
- DORA, DORA's software delivery metrics: the four keys, read 17 September 2026
- DORA, Research overview, read 17 September 2026
- Google Cloud, Announcing the 2024 DORA report, 22 October 2024
- Google, Engineering Practices: The Standard of Code Review, read 17 September 2026
- Google, Engineering Practices: Small CLs, read 17 September 2026
- Stack Overflow, 2025 Developer Survey, read 17 September 2026
- Veracode, 2025 GenAI Code Security Report, 30 July 2025
- Stripe, The Developer Coefficient, September 2018
Related reading
Why a small senior team now outbuilds a big one
Adding people used to be how you went faster. With modern tools, a small team of senior engineers often releases more work, with fewer problems, than a large mixed one.
How many tests does AI-generated code need?
The honest answer is not a number, and coverage percentages are the wrong unit. Here is the unit we use instead, and why a suite of fifteen checks can be worth more than two thousand.
More in What to assess
Assessing architecture and scalability
Architecture is the way a software system is built and organised, and scalability is whether that design keeps working under the growth the deal is paying for. Assessing both in technical due diligence means asking the team to explain the system as it is today, finding the specific parts that will break first as users, data, or traffic multiply, reading the incident history for whether root causes get fixed, and checking whether running costs rise faster than revenue. Most scaling problems are repairs with a price. The warning sign is a team that cannot answer the questions.
Assessing the engineering team
Assessing the engineering team in technical due diligence means judging whether the people can produce every future version of the software the deal is paying for. Code shows the system at one moment; the team is what turns it into the next release. The review measures seniority, the bus factor (how many people would have to leave before the project stops), tenure and turnover, how the team explains its own decisions, and how it works with AI tools. A strong senior team can fix weak code. Excellent code with too few engineers stops improving as soon as the work gets hard or a key person leaves.
Assessing security and compliance
Assessing security and compliance in technical due diligence means working out what liabilities come with the company and what it costs to reduce them. A breach or a regulatory gap can remove the value of a deal, and some of these problems take months to fix. The review does not need a penetration test to find the largest risks. It checks basic security practices (who can access production, where secrets are stored, whether systems are patched), compares the target's controls with the OWASP Top 10 and the NIST Cybersecurity Framework, reads the incident history, and confirms which certifications are current rather than planned.
Assessing technical debt
Technical debt is the extra work created by quick solutions a company chose to move fast, plus the ongoing cost they add later as slower, riskier changes. Assessing it in technical due diligence means telling acceptable debt from debt that will stop the roadmap, measuring it well enough to price, and checking whether the team has a plan to fix it. Every real codebase has debt, so the finding is never that debt exists. The finding is where it is, whether the team can name it, and whether releases are getting slower because of it.