Technical due diligence: a complete guide for investors / Special topics
SaaS metrics versus engineering reality: when the ARR and the codebase disagree
SaaS metrics and engineering reality disagree when the numbers in the deck describe a business the software cannot deliver: an ARR figure that includes services revenue the platform does not produce, a gross margin that leaves out the cloud bill, a churn rate that the incident log contradicts, or a roadmap the delivery metrics say will take three times as long. Technical due diligence is where those disagreements are found, because it is the only workstream that reads both the model and the system. This page lists the six places the two most often diverge, what artefact settles each one, and how to write the difference up as a finding the deal can price.
Published July 27, 2026. Updated September 17, 2026. Editorial.
Key takeaways
- Every SaaS metric in a deck makes a claim about the software, and each claim can be checked against an engineering artefact: the invoices, the incident log, the pipeline, the usage data.
- a16z's definition says ARR should exclude one-time fees and professional services, and the first check is whether the target's ARR does.
- Gross margin without the full cloud and inference bill is the most common disagreement, and the invoices settle it.
- The SEC's 2020 guidance on metrics asks for a clear definition, how it is calculated and how management uses it, which is the same test an investor should apply to a private company's numbers.
SaaS metrics and engineering reality disagree in six predictable places, and technical due diligence is the workstream that finds them because it is the only one with the code, the invoices and the model open at the same time. The financial reviewer checks that the ARR adds up. The technical reviewer checks that the software can produce it.
The disagreements are rarely fraud. They are usually definitions stretched in the target's favour, costs left out of a margin, and a roadmap written by someone who did not ask the engineers. Reveneau's review reconciles each metric in the deck against the engineering artefact that supports or contradicts it, and reports the gap as a finding with a cost, because a metric that the system cannot deliver is a repair bill or a price adjustment, and the deal team needs to know which.
This page assumes the reader has the pillar guide and what technical due diligence is in hand. For the cloud bill in detail, see cloud cost due diligence.
What does a SaaS metric claim about the software?
Every SaaS metric makes a claim about the software that can be checked, and the table below pairs each common metric with the engineering artefact that checks it. The definitions come from a16z's 16 Startup Metrics (2015-08-21), which remains the reference most investors use.
| Metric in the deck | What it claims about the software | Artefact that checks it |
|---|---|---|
| ARR | Revenue that recurs because the platform delivers the service without new work each period | Contract list against the product: which lines are subscriptions the software fulfils, and which are services people deliver |
| Gross margin | The cost of delivering the service, including hosting and inference, is what the model says | Twelve months of cloud and model provider invoices |
| Net revenue retention and churn | Customers stay and expand because the product works | The incident log, the support ticket volume, and usage per account over time |
| Bookings and pipeline | Contracts signed can be delivered by the product as it exists | The list of features promised in contracts against the roadmap and the delivery metrics |
| Roadmap and product velocity | The team ships at the pace the plan needs | DORA delivery metrics from the pipeline, and the last two quarters of shipped work against what was planned |
| Customer count and usage | The platform serves the number of tenants claimed at the usage claimed | Production usage data, tenant count in the database, and the cloud bill's per-unit trend |
Each row is a reconciliation. Where the deck and the artefact agree, the metric is supported. Where they disagree, the size of the gap is the finding.
How should ARR be checked against the platform?
ARR should be checked by asking which revenue lines the software produces on its own and which need people, because a16z's definition is that ARR "should exclude one-time (non-recurring) fees and professional service fees", and targets stretch it. The same guide separates bookings ("the value of a contract between the company and the customer") from revenue ("recognized when the service is actually provided"), and warns that "letters of intent and verbal agreements are neither revenue nor bookings".
The engineering check is concrete. Take the largest twenty contracts and, for each, ask what the product does for that customer without human work. Implementation fees, custom integration work, managed services and dedicated support are revenue that people deliver, and they come out of ARR. A company whose ARR is a third services is a consultancy with a software product, and the plan to scale it as SaaS assumes engineering work that has not been done: turning the services into product. That work has a cost, and it goes in the report.
Watch also for custom builds. A contract whose value depends on a feature that exists only for that customer, maintained by hand, is recurring for as long as someone maintains it. The codebase shows these as per-customer branches, configuration files with customer names, or a services directory with one module per account. Each is ARR with an engineer attached.
Why does gross margin disagree with the cloud bill?
Gross margin disagrees with the cloud bill when the margin leaves costs out, and the invoices are the only place to find them. a16z's guidance is that "all costs associated with the manufacturing, delivery, and support of a product/service should be included" and that a founder should "be prepared to break down what's included in and excluded from that gross profit figure". The reviewer's job is to do the breakdown from the source.
Rebuild the margin: revenue from the contracts, minus hosting from the cloud invoices, minus model inference from the provider invoices, minus third-party services the product calls per transaction, minus the support and customer-success people whose work delivers the service. Compare the result with the deck. A gap of a few points is a definitional difference to note. A gap of twenty points is a finding.
The AI case is where the gap is largest. Bessemer's The State of AI 2025 (2025-08-13) reported that its fastest-growing AI companies had "only 25% gross margins", and ICONIQ's State of AI 2026 (July 2026, surveys of over 300 executives) put the average at 45 percent in 2025 with projections of 53 percent in 2026 and 59 percent in 2027. Those are surveys across many companies. The target's number is its own, and rebuilding inference gross margin from invoices shows the arithmetic. A deck margin above the survey averages for a company at the same stage is a claim to check line by line.
What does the incident log say about churn?
The incident log says whether the retention figure is a property of the product or a property of the sales team, because customers who suffer outages leave, and the log records the outages. a16z defines gross churn as "MRR lost in a given month/MRR at the beginning of the month" and net churn as the same after upsells, and both are financial measures of something the engineering artefacts explain.
Lay the twelve-month incident log against the twelve-month churn series. A cluster of churn two months after a major outage is a product problem with a date on it. Churn that is flat while incidents rise means either a captive customer base or a lag that has not arrived yet. Then look at usage per account: a customer whose usage has fallen to zero and whose contract has not yet ended is churn the metric has not recorded, and a company with many of these has a retention figure that is a year out of date.
Support ticket volume is the third artefact. Rising tickets per customer with flat headcount means either the product is getting harder to use or the team is falling behind, and both show up in the retention figure eventually. Assessing the engineering team covers what falling behind looks like from the inside.
Can the team deliver the roadmap the plan assumes?
The team can deliver the roadmap when the delivery metrics support the pace the plan needs, and the pipeline is the artefact that shows it. DORA's four keys guide defines change lead time, deployment frequency, failed deployment recovery time and change fail rate as the measures of software delivery performance, and a target that can pull them from its own pipeline gives the reviewer a measured pace.
Take the roadmap in the deck and the last two quarters of shipped work. Ask what was planned two quarters ago and what shipped. The ratio is the team's real velocity, and applying it to the next four quarters of roadmap gives a delivery date the plan can be checked against. A roadmap that assumes three times the demonstrated pace is a hiring plan in disguise, and hiring has a cost, a lead time and a risk that belong in the findings.
Then take the contractual commitments: features promised to customers in writing that are not yet built. These are engineering work with a legal deadline, and they come before anything on the roadmap. A company with a quarter of committed work and a roadmap that starts from zero has a quarter it has not told you about.
How should the disagreement be reported?
Report each disagreement as a finding with the metric, the artefact, the gap and the cost, in the four-field format from the report template. The SEC's Commission Guidance on MD&A (Release 33-10751, effective 2020-02-25) says a public company presenting a metric should give "a clear definition of the metric and how it is calculated", "the reasons why the metric provides useful information to investors", and "how management uses the metric". A private target owes its investors the same three things, and a metric the target cannot define is the finding before any reconciliation starts.
Three outcomes cover most cases. The metric is supported by the artefact: say so, because supported metrics are evidence for the deal. The metric is overstated by a definitional stretch: restate it with the tight definition and give both numbers. The metric depends on engineering work not yet done: cost the work, and hand the deal team a choice between a price adjustment and a funded plan. The point of finding the disagreement early is that each of those is cheaper before signing than after, and a finding with a cost attached is one the deal can act on.
Best for
- Growth and buyout deals where the price is a multiple of ARR
- Investors who have financial diligence and want the technical review to reconcile against it
- AI companies where the presented gross margin exceeds the survey averages
Avoid if
- You want financial diligence itself, which audits the numbers rather than the system behind them
- The company is pre-revenue and the deck has no metrics to reconcile
Verify before you commit
- ARR against the top twenty contracts, line by line, for services and one-time fees
- Gross margin against twelve months of cloud and model provider invoices
- Roadmap pace against two quarters of planned versus shipped work from the pipeline
Common questions
What does it mean when SaaS metrics and engineering reality disagree?
SaaS metrics and engineering reality disagree when the numbers in the deck describe a business the software cannot deliver: ARR that includes services revenue, a gross margin that leaves out the cloud bill, churn the incident log contradicts, or a roadmap the delivery metrics say is three times too fast. Technical due diligence finds these because it reads the model and the system together.
What should ARR exclude according to the standard definition?
a16z's 16 Startup Metrics (2015-08-21) defines ARR as a measure of revenue components that are recurring in nature and says it should exclude one-time fees and professional service fees. The same guide separates bookings, the value of a signed contract, from revenue, recognised when the service is provided, and states that letters of intent and verbal agreements are neither.
How do you check ARR against the codebase?
Take the largest twenty contracts and ask, for each, what the product does for that customer without human work. Implementation, custom integration, managed services and dedicated support are revenue that the platform does not produce on its own. Per-customer branches, configuration files with customer names and one-module-per-account directories in the code are ARR with an engineer attached to it.
Why does gross margin in a deck often differ from the cloud bill?
Gross margin differs from the cloud bill when costs are left out. a16z's guidance is that all costs of delivering and supporting the service belong in gross profit and that founders should be ready to break down what is included. Rebuilding the margin from twelve months of cloud, inference and per-transaction service invoices, plus the people who deliver the service, produces the number to compare.
What gross margins do AI companies actually report?
Bessemer's The State of AI 2025 (2025-08-13) reported that the fastest-growing AI companies it studied had only 25 percent gross margins. ICONIQ's State of AI 2026 (July 2026), from surveys of over 300 executives, put average gross margin at 45 percent in 2025, with projections of 53 percent in 2026 and 59 percent in 2027. A deck margin above those for a company at the same stage is a claim to rebuild from invoices.
How does the incident log relate to churn?
The incident log explains churn because customers who suffer outages leave, and the log records the outages with dates. Laying the twelve-month incident log against the churn series shows whether churn clusters after major incidents. a16z defines gross churn as MRR lost in a month divided by MRR at the start of it, and the log is the engineering artefact behind that financial figure.
How can a reviewer tell whether the team can deliver the roadmap?
Compare what was planned two quarters ago with what shipped, and apply that ratio to the next four quarters of roadmap. DORA's four keys guide defines change lead time, deployment frequency, failed deployment recovery time and change fail rate as the measures of delivery performance, and a pipeline that produces them gives a measured pace. A roadmap assuming three times the demonstrated pace is a hiring plan.
What are contractual feature commitments and why do they matter?
Contractual feature commitments are features promised to customers in writing that are not yet built. They are engineering work with a legal deadline, so they come before anything on the roadmap. A target with a quarter of committed work and a roadmap that starts from zero has a quarter it has not disclosed, and the questionnaire on this site asks for the list directly.
What standard should a private company's metrics meet?
The SEC's Commission Guidance on MD&A (Release 33-10751, effective 2020-02-25) expects a public company presenting a metric to give a clear definition and calculation, the reasons it is useful to investors, and how management uses it. A private target owes its investors the same three things, and a metric the target cannot define is a finding before any reconciliation begins.
How should a metric disagreement be written up in the diligence report?
Write each disagreement as a finding with the metric, the artefact that checks it, the size of the gap and the cost, in the four-field format the report template uses. Three outcomes cover most cases: the metric is supported, the metric is overstated by a definitional stretch and is restated, or the metric depends on engineering work not yet done, which is costed and handed to the deal team.
References
- a16z, 16 Startup Metrics, 2015-08-21
- SEC, Commission Guidance on Management's Discussion and Analysis of Financial Condition and Results of Operations, Release 33-10751, effective 2020-02-25
- Bessemer Venture Partners, The State of AI 2025, 2025-08-13
- ICONIQ, State of AI 2026: The Builder's Economy, July 2026
- DORA, DORA's software delivery metrics: the four keys, read 2026-09-17
Related reading
Technical due diligence for VC portfolio companies
Before you commit capital, you need a clear read on the code, the team, and the risk behind it. Here is what a real technical due diligence review covers.
An AI demo is not a product
A convincing AI demo takes an afternoon. Turning it into something people trust in production is where most of the work, and most of the failures, live.
More in Special topics
Engineering artefacts for the data room: what goes in which folder
The engineering artefacts an investor should ask for in the data room are the documents and access grants that let a reviewer verify the technology instead of hearing about it: the repository inventory, the architecture as it exists, the incident log, the delivery metrics, the security reports, the open source inventory, the cloud invoices, and the engineer roster. Standard data room guidance covers none of this. a16z's guide to data rooms lists five categories to include and five to leave out, and engineering appears in neither list. This page is written for the investor requesting the room; the founder assembling it has a separate page in the preparation guide.
Open source licence due diligence
Open source licence due diligence is the part of a technical review that finds out which open source components a company's software contains, what each component's licence requires in return, and whether any of those requirements attach to the company's own code. The risk with a name is copyleft: a licence that says anyone who distributes a modified version must offer the source of the whole work under the same terms. Morgan Lewis's June 2026 note on M&A diligence lists the written policy, the inventory, the scans and the copyleft approval process as the things buyers ask about, and this page explains each in plain words for an investor who is not a lawyer.
Cloud cost due diligence: reading the bill before the deal
Cloud cost due diligence is the part of a technical review that reads the target's cloud invoices for the last twelve months, works out what the system costs to run per unit of usage, and finds the commitments, the waste and the growth curve that the deck does not show. The bill is the one engineering artefact nobody can argue with: it is what the provider charged. Flexera's 2026 State of the Cloud survey of more than 750 cloud decision-makers estimated 29 percent of cloud spend as waste, and a target's share of that is a repair bill the buyer can price before close instead of discovering after.