The money read

Rebuilding inference gross margin from model-provider invoices

To rebuild inference gross margin, take three documents for the same twelve months, the model-provider invoices, the usage logs that explain them, and revenue by customer, and compute cost against revenue per customer per month rather than accepting the blended figure in the deck. The benchmarks are wide: Bessemer's August 2025 data puts the fastest-growing AI companies at 25 percent gross margin and often negative against 60 percent for the steadier group, and ICONIQ's July 2026 survey of over 300 executives puts the 2025 average at 45 percent. This page is the worked method, using the providers' published list prices as the unit costs.

Published September 17, 2026. Editorial.

Key takeaways

  • Rebuild margin per customer per month from invoices, usage logs and revenue; a blended figure hides the contracts where usage costs more than the customer pays.
  • Use the providers' own list prices as unit costs: Anthropic lists Claude Sonnet 5 at $2 per million input tokens and $10 output, and OpenAI lists GPT-5.6 Terra at $2.00 input and $12.00 output, each with cache and batch discounts on the same pages.
  • Bessemer's August 2025 benchmarks are 25 percent gross margin, often negative, for Supernovas and 60 percent for Shooting Stars; ICONIQ's July 2026 survey reports 45 percent for 2025 with respondents projecting 53 for 2026 and 59 for 2027, and the projections are hopes rather than results.
  • Three levers move the number more than any pricing change: prompt caching at a tenth of the input price, batch processing at half price, and a model swap verified by the eval suite.

You rebuild inference gross margin by taking three documents for the same period, the model-provider invoices, the usage logs that explain each invoice line, and revenue by customer, and computing cost against revenue per customer per month. The output is a table, and the table is the finding. A blended margin from the deck tells you what the company earns on average. The table tells you which customers it loses money on, and by how much.

This page is the worked method. It uses the providers' published list prices as unit costs, and it uses invented volumes in the example so that the arithmetic is visible. The general treatment of the hosting bill is on cloud cost due diligence; this page is about the line on that bill that scales with the customer's use of the product.

What are the benchmarks, and what do they mean?

The benchmarks say that AI gross margins are lower than software margins and spread widely, and that the published projections describe hope rather than results. Reveneau's margin rebuild ignores the target's projected margin entirely and uses only the invoices, because the invoices are the one document in the data room that the provider wrote rather than the company.

Bessemer's State of AI 2025, published 13 August 2025, describes two archetypes. "Supernovas" reach $40 million of first-year revenue, and Bessemer gives their gross margin as 25 percent while noting it is often negative, with low switching costs and margin compression as the defining features. "Shooting Stars" grow more slowly, to a first year Bessemer gives as $3 million, at a gross margin of 60 percent [1]. Bessemer gives both margins as approximations. ICONIQ's State of AI 2026, from a July 2026 survey of over 300 software executives, reports an average gross margin of 45 percent for 2025 and respondent projections of 53 percent for 2026 and 59 percent for 2027; it also reports that two-thirds of respondents say their per-query unit economics improved [2].

For context on what "software margin" means, a16z's June 2025 analysis by Joe Schmidt gives ServiceNow's gross margin at IPO as 63.2 percent and Workday's as 54.1 percent, both well below the 80 percent the industry treats as the benchmark, and both later reaching 79 and 75 percent in 2024 [3]. The point of that comparison is that low margin at scale has a history of recovering, and the point of this page is that you find out whether this company's margin can recover by looking at where it is being lost.

The method, step by step

The method has seven steps and the first three are requests.

  1. Ask for twelve months of model-provider invoices. Every provider: the model labs directly, and any cloud marketplace through which models are billed. The invoice is the cost of goods for inference and it is the document the company did not write.
  2. Ask for the usage logs behind them. Tokens in, tokens out, cache reads, cache writes, batch versus interactive, by model, by customer, by day. If the company cannot attribute usage to customers, that is the first finding, and the rest of the rebuild is done at the company level with a note saying so.
  3. Ask for revenue by customer by month. The same months, the same customer identifiers.
  4. Price the usage at list. Use the providers' published pages. Anthropic lists Claude Sonnet 5 at $2 per million input tokens and $10 per million output tokens, with cache reads at a tenth of the input price and a 50 percent discount for batch processing; Claude Opus 5 is $5 and $25; Claude Haiku 4.5 is $1 and $5 [4]. OpenAI lists GPT-5.6 Terra at $2.00 input, $0.20 cached input and $12.00 output per million tokens, GPT-5.6 Sol at $4.00, $0.40 and $20.00, and GPT-5.6 Luna at $0.20, $0.02 and $1.20 [5]. Reconcile the priced usage to the invoice total; a gap is a negotiated discount or a missing log, and either is worth knowing.
  5. Compute inference cost per customer per month. Then margin: revenue minus inference cost, over revenue.
  6. Add the rest of cost of goods. Hosting, third-party data, support, at the company's allocation. Inference alone is the AI-specific line; gross margin needs all of it.
  7. Sort the table by margin. The bottom rows are the finding.

A worked example

Suppose one customer pays $2,000 a month and the usage logs show 400 million input tokens and 40 million output tokens on Claude Sonnet 5 in that month. The volumes are invented; the prices are Anthropic's list prices [4].

Line Volume List price Cost
Input tokens 400 million $2 per million $800
Output tokens 40 million $10 per million $400
Inference cost $1,200
Revenue $2,000
Inference gross margin 40 percent

Now suppose the logs show that 300 million of those input tokens were cache reads, which Anthropic prices at a tenth of the input rate, $0.20 per million for Sonnet 5 [4].

Line Volume List price Cost
Uncached input 100 million $2 per million $200
Cache reads 300 million $0.20 per million $60
Output tokens 40 million $10 per million $400
Inference cost $660
Inference gross margin 67 percent

The same usage on OpenAI's GPT-5.6 Terra, with no caching, at $2.00 input and $12.00 output, costs $800 plus $480, or $1,280, a margin of 36 percent [5]. On Claude Haiku 4.5 at $1 and $5, it costs $400 plus $200, or $600, a margin of 70 percent, if the eval suite shows the cheaper model meets the specification [4]. That last condition is the whole of the eval suite as a diligence artefact: a model swap is a margin lever only if the company can prove it did not break the product.

What does the per-customer view show?

The per-customer view shows which contracts are subsidised, and it shows it in a way no blended figure can. In the example above, the company might report 55 percent blended inference margin across its base. The table might show three customers at 70 percent, ten at 50, and two large accounts at minus 20, because those two were sold a flat fee and use the product all day. The blended figure is true. The two accounts are the finding, and the question for the management meeting is what happens at renewal.

The same view answers three other questions. Whether pricing matches usage: ICONIQ's July 2026 survey reports consumption-based pricing rising from 35 percent to 42 percent of respondents in six months and outcome-based pricing from 18 to 23 percent [2], which is companies moving price towards cost, and the per-customer table shows whether this one has. Whether there is a cap on consumption per user, which is entry LLM10 on the OWASP list and is covered on security of AI-written code in diligence. And whether the reported usage metric matches the billed one, which is the AI version of the gap described in SaaS metrics versus engineering reality.

What moves the margin?

Three levers move inference margin more than any pricing change, and each one is visible in the logs.

Caching. Anthropic prices cache reads at 0.1 times the input price and OpenAI lists cached input at a tenth of the standard input rate for the GPT-5.6 models [4] [5]. A product with a long system prompt and repeated context that is not caching is paying ten times what it needs to on that share of its input.

Batch processing. Both providers list a 50 percent discount for asynchronous batch requests [4] [5]. Anything the product does that a user is not waiting for, nightly processing, bulk classification, embedding, can run at half price.

Model choice, verified. The Haiku example above is a 30-point margin change from one configuration line, and it is only real if the eval suite says the output still meets the spec.

The price list itself also moves, in both directions. Stanford's AI Index 2025 reports that the cost of querying a model performing at GPT-3.5 level fell from $20.00 per million tokens in November 2022 to $0.07 by October 2024, a more than 280-fold reduction, and that inference prices have fallen between 9 and 900 times a year depending on the task [6]. Against that, Anthropic's pricing page notes that the tokenizer used by Claude 4.7 and later produces more tokens for the same text, giving 30 percent as the figure, so a model upgrade can raise the bill for the same usage [4]. The per-customer table, rebuilt after a swap, shows which effect won. Model dependency risk covers what a price change or a retirement does to the plan.

What goes in the report?

The report contains the per-customer table sorted by margin, the reconciliation of priced usage to the invoices, the three levers and how much each is worth for this product, and one sentence on what the plan assumes about price. A company that already has the table sends it in a day and the finding is the bottom rows. A company that does not has never seen its own worst customers, and that is a finding about the finance function as well as the margin. AI feature or AI wrapper uses the same table for a different purpose, and the AI startup due diligence guide puts the margin rebuild at the centre of the money read.

Best for

  • Any AI deal where gross margin is in the model and the target's figure is blended
  • A deal team that has the invoices and a spreadsheet afternoon
  • A founder who wants to see the table investors will build

Avoid if

  • The product runs its own models on its own hardware; the cost lines are different and belong on the cloud cost page
  • Inference is under a tenth of cost of goods and the deal does not turn on it

Verify before you commit

  • Reconcile usage priced at list to the invoice totals for the same months
  • Sort the per-customer table and read the bottom five rows with the contract terms
  • Check whether the caching and batch discounts on the provider pages are being used

Common questions

How do you calculate an AI startup's gross margin from its invoices?

You calculate an AI startup's inference gross margin by pricing the usage logs at the providers' list prices, reconciling to the invoices, and dividing revenue minus inference cost by revenue for each customer and month. Anthropic lists Claude Sonnet 5 at $2 per million input tokens and $10 output, and OpenAI lists GPT-5.6 Terra at $2.00 and $12.00, with cache and batch discounts on the same pages. Then add hosting and the rest of cost of goods, and sort the table.

What is a normal gross margin for an AI startup?

A normal gross margin for an AI startup is lower than for software and spread widely. ICONIQ's July 2026 survey of over 300 software executives reports 45 percent for 2025, with projections of 53 percent for 2026 and 59 percent for 2027. Bessemer's August 2025 data splits fast growers into Supernovas at 25 percent, often negative, and Shooting Stars at 60 percent. a16z's June 2025 analysis gives 80 percent as the software benchmark. The target's figure comes from its own invoices.

Why is a blended gross margin figure misleading for an AI company?

A blended gross margin figure is misleading for an AI company because inference cost scales with each customer's usage, so an average hides the contracts where usage costs more than the customer pays. A company can report 55 percent blended while two large flat-fee accounts run at a loss all year. Bessemer's August 2025 finding that the fastest-growing AI companies run at 25 percent gross margin and often negative is what that pattern looks like at scale. Build the table per customer per month.

What documents do I need to rebuild inference margin?

You need three documents for the same twelve months: the model-provider invoices from every lab and cloud marketplace the company is billed through, the usage logs that explain each invoice line by tokens, model, customer and day, and revenue by customer by month. If usage cannot be attributed to customers, that is the first finding. Anthropic's and OpenAI's pricing pages supply the unit costs; price the usage at list and reconcile to the invoice total before computing anything.

How much does prompt caching change an AI startup's margin?

Prompt caching can change an AI startup's inference margin by tens of points when a large share of input is repeated context. Anthropic prices cache reads at a tenth of the input price, $0.20 per million tokens for Claude Sonnet 5 against $2 uncached, and OpenAI lists cached input for GPT-5.6 Terra at $0.20 against $2.00. In the worked example on this page, moving 300 million of 400 million input tokens to cache reads lifts the margin on one customer from 40 percent to 67 percent.

Can switching to a cheaper model fix a bad AI margin?

Switching to a cheaper model can fix a bad AI margin only if the eval suite shows the cheaper model still meets the specification. Anthropic lists Claude Haiku 4.5 at $1 input and $5 output against $2 and $10 for Sonnet 5, so the same usage costs half; in the worked example that is a move from 40 percent to 70 percent margin on one customer. Without an eval run on the new model, the swap is a bet on quality, and the margin gain is a bet on retention.

Are inference prices going up or down?

Inference prices have fallen sharply for a fixed level of capability and can rise for a given product when it upgrades. Stanford's AI Index 2025 reports that the cost of a model at GPT-3.5 level fell from $20.00 per million tokens in November 2022 to $0.07 by October 2024, a more than 280-fold reduction. Against that, Anthropic's pricing page states that its newer tokenizer produces 30 percent more tokens for the same text on Claude 4.7 and later, so an upgrade can raise the bill for the same usage.

What does the batch discount do for an AI startup's cost of goods?

The batch discount halves the cost of any model work a user is not waiting for. Anthropic and OpenAI both list a 50 percent discount on input and output tokens for asynchronous batch processing on their pricing pages. Nightly processing, bulk classification and embedding jobs can all run at half price, and the usage logs show whether they do. A product whose logs show every call as interactive is paying full price for work that could be queued.

Should an AI startup price by consumption?

Whether an AI startup should price by consumption depends on the product, but the per-customer margin table shows whether its current pricing matches its cost. ICONIQ's July 2026 survey reports consumption-based pricing rising from 35 to 42 percent of respondents in six months and outcome-based pricing from 18 to 23 percent, which is the industry moving price towards cost. A flat-fee contract with an unbounded user is the pattern that produces negative margin rows in the table.

What should the diligence report say about inference margin?

The diligence report should include the per-customer table sorted by margin, the reconciliation of priced usage to the invoices, the value of the three levers for this product, caching, batch and a verified model swap, and one sentence on what the plan assumes about the providers' prices. ICONIQ's July 2026 respondents project 53 percent for 2026 and 59 percent for 2027; the report should say whether this company's plan relies on the same hope and what it rests on.