Commercials

Measuring forward deployed engineering

Measure four things: time from contract signature to production, the share of engagements that reach production, revenue per embedded engineer over time, and what came back into the product. Do not measure utilisation. Utilisation rewards keeping engineers on accounts, which is exactly the endless engagement this model fails into. The metric that matters most to a customer is time to production, and the metric that tells you whether you have a business is revenue per engineer rising rather than staying flat.

Key takeaways

  • Four metrics: signature-to-production time, share of engagements reaching production, revenue per embedded engineer over time, and capabilities returned to the product.
  • Do not measure utilisation. It rewards keeping an engineer on an account, which is the failure mode of the whole model.
  • Signature-to-production should fall across successive deployments. That is the number a customer feels.
  • DORA's long-running research finds the strongest teams ship more often and break things less, so speed and stability are not a trade here either.

Most measurement of embedded engineering functions is borrowed from professional services, which is why it drives the wrong behaviour. Here is what to track instead.

The four that work

Time from contract signature to production. Not from kickoff. From signature, because the gap between signature and somebody starting is real and it is yours to fix.

This is the number a customer experiences, and it should fall across successive deployments as the platform absorbs what earlier ones learned. A flat signature-to-production time over eight deployments means nothing is being reused. It is also the number a16z's Marc Andrusko suggests investors ask about directly: how many engineer-months from contract signature to production.

Share of engagements that reach production. Count started engagements and count the ones running in production for real users. This is the honest version of a success rate, and it should be embarrassing at first.

Compare it against the context: MIT NANDA reported in July 2025 that 95 percent of generative AI pilots produced no measurable profit-and-loss impact, from 52 executive interviews, 153 leader surveys and 300 public deployments. If your rate is well above that, it is your strongest sales asset and you should be able to prove it.

Revenue per embedded engineer, over time. The business-model measure. It should rise, because the platform is absorbing custom work and each engineer covers more ground. Its own page covers the detail, and the summary is that flat revenue per engineer with growing custom work means you are running a services shop.

Capabilities returned to the product. Count them, per period, with a name attached. This is the output of the productization loop and the only direct measure of whether the margin you spent bought anything durable.

The three that look useful and are not

Utilisation. The standard professional services metric and the most damaging one here. It rewards keeping engineers on accounts, which is the endless engagement. A function at 95 percent utilisation is a function with no time to generalise, which means it is compounding nothing. If you track it at all, track it as a ceiling not a target.

Billable hours. Same problem, and it also pays the vendor to be slow. How to price covers what that does to behaviour.

Customer satisfaction alone. A customer whose engineers never learned the system and who now depends on you is frequently delighted. Satisfaction is worth knowing and it does not distinguish a successful deployment from a comfortable dependency. Pair it with whether the customer's own engineer has shipped a change unaided.

Two more worth tracking quietly

Share of shared versus custom code per deployment. Should trend toward shared. If the custom share grows with each account, the platform is losing to the field.

Handover completion rate. The share of engagements where the customer's team took the system over on the agreed date and still runs it a quarter later. This is the metric that detects the dependency failure mode, and almost nobody tracks it.

Do not trade speed for stability

Worth stating because the instinct in an embedded engagement is to slow down for safety in somebody else's production environment.

Google's DORA research, which has studied software delivery for over a decade, finds that the strongest teams ship more often and break things less at the same time. Speed and stability rise together rather than trading off. The mechanism is the verification, not caution.

In this context that means the way to move quickly inside a customer's environment is better checks rather than fewer changes. Generated code is not secure by default, so if a model is writing it, the checks have to be derived from the specification and graded independently of the model that produced the code. That is eval-driven development, and it is what makes a fast embedded engagement safe rather than reckless.

What to report, and to whom

To the board or leadership: signature-to-production time, share reaching production, revenue per engineer, capabilities returned. Four numbers, tracked as a trend rather than a snapshot.

To the customer: what shipped this week, progress against the production date, and the honest list of what has been discovered that changes the plan. Not utilisation, not hours.

To yourself, weekly: the time split per engineer against the roughly 20 percent customer-facing and 80 percent building that First Round Review describes. It is the earliest warning that an engagement is going wrong, and it moves before any of the outcome metrics do.

The one-line test

If your metrics would look better on an engagement that never ended, they are the wrong metrics.

Every measure on the recommended list gets worse when an engagement drags: signature-to-production rises, revenue per engineer flattens, handover completion falls. Every measure on the discouraged list gets better. That is the whole reason to choose between them deliberately.

Best for

  • Choosing the metrics for a new embedded engineering function
  • Replacing professional services metrics inherited from a services org
  • Reporting the function's value to leadership

Avoid if

  • You need the pricing model rather than the measurement set

Verify before you commit

  • Confirm utilisation is not a target, and ideally not tracked at all
  • Measure signature-to-production rather than kickoff-to-production
  • Track handover completion a quarter after the handover date
  • Apply the test: would these metrics look better on an engagement that never ended

Common questions

What should you measure in a forward deployed engineering function?

Four things. Time from contract signature to production, the share of engagements that reach production, revenue per embedded engineer tracked over time, and the number of capabilities returned to the product with names attached. Track all four as trends rather than snapshots.

Why should you not measure utilisation?

Because it rewards keeping engineers on accounts, which is the endless engagement this model fails into. A function at 95 percent utilisation has no time to generalise, so it is compounding nothing. If you track it at all, treat it as a ceiling rather than a target.

Why measure from signature rather than kickoff?

Because the gap between a contract being signed and somebody actually starting is real, it is yours to fix, and measuring from kickoff hides it. Signature-to-production is also the number a16z suggests investors ask about directly, and it is what the customer experiences.

Is customer satisfaction a good measure?

On its own, no. A customer whose engineers never learned the system and who now depends on you is frequently delighted. Satisfaction does not distinguish a successful deployment from a comfortable dependency, so pair it with whether the customer's own engineer has shipped a change unaided.

Should you slow down to be safe in a customer's production environment?

The evidence says no. Google's DORA research finds the strongest teams ship more often and break things less at the same time, so speed and stability rise together. The mechanism is verification rather than caution, which means better checks rather than fewer changes.

What is the simplest test of whether your metrics are right?

Ask whether they would look better on an engagement that never ended. Signature-to-production, revenue per engineer and handover completion all get worse when an engagement drags. Utilisation and billable hours get better. That difference is the whole reason to choose deliberately.

More in Commercials

How to price forward deployed engineering

There are four ways to price embedded engineering: time and materials, fixed scope, outcome, and bundled into a licence. Each produces different behaviour, and the behaviour matters more than the rate. Billable hours pay the vendor to be slow and to keep the engagement going. Fixed scope pays them to argue about scope. Outcome pricing aligns both sides and is hard to write. When AWS announced its own forward deployed engineering organization it said the work would be structured around shared goals and business results rather than billable hours, which is a standard worth holding every vendor to.

Outcome-based pricing for embedded engineering

An outcome you cannot measure on the day of signature is a hope with a payment schedule attached. A workable outcome passes three tests: it is observable rather than inferred, somebody named takes the measurement, and neither party can move the goalposts unilaterally. The most reliable outcome for an embedded engagement is the least exciting one, which is a defined system running in production for real users by a named date. Business-metric outcomes sound better and fail more often, because neither side fully controls them.

What forward deployed engineering costs

Two compensation figures for this role are traceable to a source. Levels.fyi reports Palantir forward deployed engineer total compensation in the United States with a median of $278,000 and a range from $185,000 to $631,000, read on 8 September 2026. Across 1,000 job postings analysed by Revealera and published in November 2025, the median advertised salary was $173,816. Everything else circulating is unsourced. Load one of those two for benefits, travel and unbilled time to get your input cost, then compare it against the cost of the deployment not happening.

Revenue per forward deployed engineer

Revenue per embedded engineer, tracked over time, is the single number that tells you which business you are in. It has to rise. When the platform absorbs what the field learned, each engineer covers more ground and the number climbs. If it stays flat while custom work accumulates, you are running a services shop that has not admitted it yet. The benchmark for what a rising curve looks like is the previous generation: ServiceNow's gross margin at IPO was 63.2 percent and Workday's 54.1 percent, reaching 79 and 75 percent by 2024.

The contract for an embedded engagement

Before an outside engineer touches your repository, six things belong in writing: IP assignment on creation, scoped repository access rather than organization-wide, production access as a named exception with an audit trail, a production date, a handover date with a defined artefact, and subcontractor binding. The one people get wrong most often is the first: an NDA covers secrecy and does not assign ownership of what the vendor creates. Those are two separate clauses and you need both.