AI startup due diligence: how investors check what is real / The money read
Data moat claims: what counts and what does not
A data moat claim counts when the company can show four things: that the next unit of data makes the product measurably better, that a competitor could not get equivalent data at a similar cost, that the company has the rights to use the data the way it does, and that the loop closes, meaning use of the product produces data that improves the product. Most claims fail the first test. a16z's 2019 analysis by Casado and Lauten argued that there is generally no inherent network effect from having more data, that the value of the next unit falls while its cost rises, and that a minimum viable corpus is enough to start training against and gives no lasting defence. This page gives the four tests and the evidence for each.
Published September 17, 2026. Editorial.
Key takeaways
- Ask for the curve: product quality against training or retrieval data volume. A flat curve means the company has a data asset and no moat; a rising curve means the next customer makes the product better for the last one.
- a16z's 2019 analysis is the reference argument: data has scale effects rather than network effects, the marginal unit is worth less and costs more, and reaching a minimum viable corpus is a starting line rather than a defence.
- Foundation models trained on the public internet have raised the bar for what counts as proprietary: the surviving data advantage, in Foundation Capital's September 2026 framing, is a feedback loop inside a workflow the provider does not see.
- Rights are part of the moat: the provider's terms decide whether the company's data loop stays its own, and Anthropic's commercial terms state the customer retains rights to inputs and owns outputs and that Anthropic may not train on customer content.
A data moat claim counts when four things are true and can be shown: the next unit of data makes the product measurably better; a competitor could not get equivalent data at a similar cost; the company has the rights to use the data as it does; and the loop closes, so that use of the product produces the data that improves it. A claim that passes all four is rare. A claim that passes none is the default, because every AI deck says "proprietary data" and the slide costs nothing.
This page gives the four tests, the evidence for each, and what to write when the claim fails, because a data asset with no moat is still an asset, and the report should say what it is worth.
What does the a16z argument say?
The a16z argument says that what founders call data network effects are usually data scale effects, and that scale effects are a weak defence because the value of additional data falls as the corpus grows while the cost of acquiring it rises. Reveneau's diligence asks for the quality-against-data curve before any discussion of the moat, because the curve is the whole argument in one chart and the company either has it or has never measured the thing it is claiming.
Martin Casado and Peter Lauten published "The Empty Promise of Data Moats" on 9 May 2019. Its central claims, quoted: "There generally isn't an inherent network effect that comes from merely having more data." On diminishing returns, using a chatbot example, "after 40% of queries collected, there is actually no advantage to collecting more data at all." On cost, "getting the next piece of data tends to become more expensive to capture over time," so that "the cost of adding unique data to your corpus may actually go up, while the value of incremental data goes down." And on the starting line: bootstrapping a "minimum viable corpus" is sufficient to start training against, and reaching it "clearly will not be a durable moat" [1].
Seven years on, the argument is stronger, because the models most startups build on were trained on most of the public internet. A corpus that would have been proprietary in 2019 is now a subset of what the provider already has. Foundation Capital's September 2026 analysis of model providers moving into applications makes the updated case: providers have a view across everything built on their platforms, and the data advantage that survives is a feedback loop captured inside a workflow with high exception rates, subjective success criteria and fragmented legacy systems, which is to say inside something the provider does not see and will not build [2].
What are the four tests?
The four tests are the curve, the cost of equivalent data, the rights, and the loop. Each has a specific piece of evidence, and the table gives the claim, the evidence, and what a failure means.
| Test | Ask for | Passes when | Fails when |
|---|---|---|---|
| The curve | Product quality (eval score) plotted against data volume, from the company's own runs | Quality rises with volume across the range the company operates in | Quality is flat, or rose early and plateaued at a volume a competitor could reach |
| Cost of equivalent data | The source of the data and what it cost to acquire, per unit | A competitor would need years, a licence nobody else has, or a customer base to reproduce it | The data is public, purchasable, or generated by a process anyone can run |
| Rights | The terms under which the data was collected and the provider terms it flows through | Customers agreed to the use, and the provider may not train on it | Consumer terms, silent contracts, or data the company does not own |
| The loop | A diagram of how product use produces data that changes the product, and a dated example | Use produces labelled outcomes that the evals confirm improved the next version | Data is collected and stored; nothing downstream reads it |
The curve is the test that most claims fail, and the reason is simple: measuring it requires an eval suite, because "better" has to be a number. A company without a suite cannot show the curve. It can only show the volume. The eval suite as a diligence artefact is the page on reading a suite, and the curve is the first thing to ask a suite to produce.
Why does the loop matter more than the corpus?
The loop matters more than the corpus because a corpus is a stock that a competitor can match and a loop is a flow that a competitor has to earn customer by customer. Casado and Lauten's point about the minimum viable corpus is that the stock stops mattering once it is large enough to train on [1]. What keeps mattering is whether each new customer's use of the product produces data the last customer benefits from, and whether that data is labelled by something the world would otherwise not label: a decision made, a correction entered, an outcome recorded.
The loop also has to stay the company's own. Data flowing through a model provider's API is subject to that provider's terms. Anthropic's commercial terms state that the customer "retains all rights to its Inputs" and "owns its Outputs", and that "Anthropic may not train models on Customer Content from Services" [3]. Terms differ between consumer and commercial products and between providers, so a company claiming a proprietary loop should be able to show that every hop of the loop runs under terms that keep it proprietary. Code provenance and licence exposure covers the same question for the code; the answer for the data goes in the same table.
Bessemer's August 2025 description of the fastest-growing AI companies, explosive adoption with low switching costs and margin compression [4], is what a data claim without a loop looks like from the outside: growth that a competitor with the same model and the same public data can take back. A working loop shows up in the same place a moat always shows up, which is retention. If churned customers replaced the product easily, the data was not making it better for them.
How do you test a data moat in diligence?
Test a data moat by asking for the curve first and then walking the other three tests with the evidence in hand.
- Ask for the curve. Eval score against data volume, from the company's own runs, across the volumes it has actually operated at. If there is no eval suite, record "curve unavailable" and move on; the claim is untested by definition.
- Ask where the data came from and what it cost. Per unit, with the source. Then ask what the same data would cost a well-funded competitor starting today, and whether the answer is money or time.
- Read the rights. Customer contracts for the collection, provider terms for the processing, account types and dates. A period on consumer terms is a period in which the loop may have leaked.
- Trace one loop, dated. One example of product use producing data that changed a later version, with the eval run that measured the change. A company with a loop has this in its release history.
- Ask what the provider sees. Which parts of the loop run through a third-party model, and whether the provider could reproduce the loop from what passes through its API. The claim test on how to verify an AI claim can include an input designed to see whether the company's data changes the answer at all.
- Read the churn notes. Where did lost customers go, and did they cite quality. A moat that does not show up in retention is a slide.
What to write when the claim fails
Write what the data is worth without the moat. A labelled corpus that a competitor could match in a year is still a year's lead, and a year is worth something at the right price. Training data that was expensive to collect and is now a subset of what the model provider has is worth its licence value and little more. A collection pipeline that produces useful data but feeds nothing is an asset waiting for an engineering project, and the project belongs in the post-close plan.
The one-line report reads: "Data claim: curve unavailable (no eval suite); corpus purchasable at a cost the company could not state; rights clean on commercial terms from March; loop open. Value the data as a lead." Or, when it passes: "Data claim: curve rises across operating range per the suite; source is customer outcomes unavailable elsewhere; rights clean; one dated loop traced in the June release. Value as a moat and verify at renewal data."
AI feature or AI wrapper is the page for the product-side version of the same question, model dependency risk covers what a loop is worth on the day the provider ships the product, and the AI startup due diligence guide puts the data question at the end of the money read because its answer depends on everything before it. Execution speed is the moat is our own view on what defends an AI product when the data does not.
Best for
- Any AI deal where 'proprietary data' appears on a slide and in the valuation
- A deal team deciding how much of the price rests on the moat
- A founder who wants to know what evidence turns the slide into a claim
Avoid if
- The company makes no data claim and competes on workflow or distribution
- You have no eval suite to produce the curve; run the other three tests and mark the first untested
Verify before you commit
- Ask for the quality-against-data curve from the company's own eval runs
- Price what equivalent data would cost a competitor, in money and in time
- Trace one dated loop from product use to a measured improvement in a later release
Common questions
What is a data moat in an AI startup?
A data moat in an AI startup is a data advantage that makes the product measurably better with each unit of additional data, that a competitor could not reproduce at a similar cost, that the company has the rights to use, and that closes into a loop where product use produces the data that improves the product. a16z's 2019 analysis by Casado and Lauten argued most data advantages are scale effects, with the marginal unit worth less and costing more, rather than moats. The curve is the first test.
Why is 'we have proprietary data' usually a weak claim?
'We have proprietary data' is usually a weak claim because having data and having a moat are different things: the value of the next unit of data falls as the corpus grows, the cost of acquiring it rises, and, in Casado and Lauten's 2019 words at a16z, reaching a minimum viable corpus is sufficient to start and will not be a durable moat. Foundation models trained on the public internet have since raised the bar further. The claim becomes strong when the company can show a rising quality-against-data curve.
What is the quality-against-data curve and why ask for it?
The quality-against-data curve plots the product's eval score against the volume of training or retrieval data, from the company's own runs, across the volumes it has operated at. It is the whole moat argument in one chart: a rising curve means more data keeps making the product better, and a flat one means the company has an asset without a defence. a16z's 2019 chatbot example found no advantage from collecting more data after 40 percent of queries; ask where this company's curve flattens.
How do I check whether a competitor could get the same data?
Check whether a competitor could get the same data by asking where the company's data came from, what it cost per unit, and what a well-funded competitor starting today would need: money, time, a licence nobody else holds, or a customer base. Public, purchasable or process-generated data fails the test. Casado and Lauten's 2019 analysis notes that the cost of adding unique data rises over time, which cuts both ways: it protects an incumbent's late data and makes its early data cheap to match.
What is a data loop and how is it different from a data corpus?
A data corpus is a stock of data the company holds; a data loop is a flow in which use of the product produces data that improves the product for the next user. The stock stops mattering once it is large enough to train on, which is the minimum viable corpus point in a16z's 2019 analysis; the flow keeps mattering because a competitor has to earn it customer by customer. Ask for one dated example of the loop closing, with the eval run that measured the improvement.
Do the model provider's terms affect a data moat?
Yes, the model provider's terms affect a data moat, because data that flows through a provider's API is subject to the provider's terms, and a loop only stays proprietary if every hop keeps it so. Anthropic's commercial terms state that the customer retains all rights to its inputs and owns its outputs, and that Anthropic may not train models on customer content from the services. Consumer terms differ, so a period on consumer accounts is a period in which the loop may have leaked.
Can the model provider reproduce a startup's data advantage?
The model provider can reproduce a startup's data advantage when the loop runs through its API in a form it can see and the workflow is one it would build. Foundation Capital's September 2026 analysis argues providers have a view across everything built on them and names the surviving advantages as feedback loops inside high-exception workflows, subjective success criteria and fragmented legacy systems. Ask which parts of the loop the provider sees and whether it could rebuild the loop from them.
How does a real data moat show up in the numbers?
A real data moat shows up in retention and in the eval history: churned customers cite something other than quality, and each release that used more data scored higher on the suite. Bessemer's August 2025 description of the fastest-growing AI companies, explosive adoption with low switching costs and margin compression, is what a data claim without a moat looks like from outside. Read the churn notes for where lost customers went and whether a competitor on the same model replaced the product easily.
What should the report say if the data moat claim fails?
If the data moat claim fails, the report should say what the data is worth without the moat: a labelled corpus a competitor could match in a year is a year's lead, an expensive training set that is now a subset of what the provider holds is worth its licence value, and a pipeline that feeds nothing is an engineering project for the post-close plan. Casado and Lauten's 2019 advice at a16z was to avoid treating data as a magical moat; the report should price it as an asset instead.
Can a startup have a data moat without an eval suite?
A startup can have a data advantage without an eval suite, but it cannot show a data moat, because the first test is the quality-against-data curve and the curve needs a number for quality. Without a suite, record the curve as unavailable, run the other three tests, cost, rights and loop, and mark the claim untested rather than false. Anthropic's developer guidance describes suites with many automated cases; building one is the first step to proving the moat as well as the product.
References
- a16z, Martin Casado and Peter Lauten, The Empty Promise of Data Moats, 9 May 2019
- Foundation Capital, When model providers eat everything: a survival guide for service-as-software startups, September 2026
- Anthropic, Commercial Terms of Service, read 17 September 2026
- Bessemer Venture Partners, The State of AI 2025, 13 August 2025
- Anthropic, Develop test cases (Claude Platform docs), read 17 September 2026
Related reading
Speed of execution is the moat now
Being first used to be a defensible lead. When any capable team can build the same thing in a week, the advantage moves to how fast you ship, learn, and ship again.
An AI demo is not a product
A convincing AI demo takes an afternoon. Turning it into something people trust in production is where most of the work, and most of the failures, live.
More in The money read
Rebuilding inference gross margin from model-provider invoices
To rebuild inference gross margin, take three documents for the same twelve months, the model-provider invoices, the usage logs that explain them, and revenue by customer, and compute cost against revenue per customer per month rather than accepting the blended figure in the deck. The benchmarks are wide: Bessemer's August 2025 data puts the fastest-growing AI companies at 25 percent gross margin and often negative against 60 percent for the steadier group, and ICONIQ's July 2026 survey of over 300 executives puts the 2025 average at 45 percent. This page is the worked method, using the providers' published list prices as the unit costs.
Model dependency: what a model swap or a price change does to the plan
Model dependency risk is the exposure an AI startup carries because the model in its product is a supplier's product, which the supplier can retire, reprice, change or compete with. The providers publish their own rules: OpenAI gives at least 6 months' notice before retiring a generally available model and Anthropic at least 60 days, and both retired models in 2026. A price change moves the margin directly, a retirement forces an unplanned migration, and a provider feature launch can replace the product. This page sets out each event, what it does to the plan, and how to test in diligence whether the company could survive it.