Adopting the practice

How to move an existing engineering team to AI-native development

Most teams adopt AI coding tools by buying licences and hoping. A few months later nobody can say whether it helped, because nothing was measured before the change and the only evidence is how people feel. This is a plan for doing it in an order that produces an answer: what to measure first, which single flow to change, why the automated checks have to come before the volume, and the two failure modes that cause most rollouts to stop.

Published July 28, 2026. Editorial.

Key takeaways

  • Measure delivery before you start, because self-reported speed is demonstrably unreliable.
  • Change one flow in one repository first, not every team at once.
  • Put the blocking checks in before you raise the volume of generated code, never after.
  • The two failure modes are a review queue that becomes the slowest step, and a passing build that proves nothing.
  • Rollouts stop for lack of evidence far more often than for lack of enthusiasm.

Rolling out AI coding tools looks like a tooling change and behaves like a process change. That mismatch is why so many attempts produce a lot of activity and no defensible conclusion.

This page is about an existing team with a working codebase and delivery commitments. Not a new build from nothing, where you can set the practice on day one.

Why rollouts stop

The common story is resistance. That is rarely what actually happens. Engineers try the tools readily enough. What goes missing is evidence.

Three months in, someone asks whether it worked. There is no baseline, so the answer is a survey of impressions. And impressions are the one measure we know is wrong here. In a randomised controlled trial published in July 2025, METR gave 16 experienced open-source developers 246 real tasks in repositories they already knew well. The developers using AI tools took 19 percent longer, and estimated afterwards that AI had made them 20 percent faster [1]. METR labels the finding historical, since the tools have changed since then, and that is fair. The part that still holds is that skilled people misjudged their own performance by 39 points, in the direction that made the tool look better.

So a rollout that plans to evaluate itself by asking the team is a rollout that has already decided its answer.

How this shows up in practice

A payments team at a mid-size software company adopted an AI coding tool for every engineer on a Monday, with no baseline and no target flow chosen. By week six, two things were both true at once and nobody could say which mattered more: the team had merged more pull requests than any six-week period in its history, and a billing calculation bug had reached production that the existing test suite should have caught. The engineering lead's report to their own manager said adoption was "going well," based on the pull request count, and did not mention the bug, because the two facts lived in different systems and nobody had been asked to look at them together. Three months later, when a budget review asked whether the licence spend was worth renewing, the only available answer was a survey where most engineers said they felt faster. That is the outcome step one and step two exist to prevent: a real change in the codebase, a real incident, and a review that ran entirely on gut feeling because delivery data was never captured going in.

Step one: measure before you change anything

Two weeks of baseline is enough, and four metrics cover it. How long a change takes from first commit to production. How often you deploy. What share of changes cause a problem that needs a fix. How long recovery takes when one does. These are DORA's four keys and they are well understood [2].

Pull them from the tools you already have. Do not build a dashboard, and do not add a metric about volume. Lines of code, number of pull requests, and commits per developer all go up with AI adoption whether or not anything improved, which makes them harmful: they will show a win in every scenario, including the bad ones.

Step two: one flow, one repository

Pick a single service with real users, an existing test suite, and one team that owns it. Not the most critical system, and not a practice project.

Change one flow inside it: the path from a written ticket to a merged change. That means the specification gets written to the level of detail a model can build against, the implementation is generated, and the automated checks decide whether it merges. Everything else about how the team works stays the same.

The instinct is to roll out to everyone at once, since the licence is per seat and enthusiasm is high. Resist it. A change in one place produces a comparison. The same change everywhere produces only impressions.

An edge case worth planning for before it happens: the target service turns out to have a test suite that runs, but that nobody has broken on purpose in over a year. That suite might be asserting real behaviour, or it might be asserting nothing and simply never failing because nobody has tried to make it fail. Treat this as a finding, not a delay. Spend a day breaking three or four of its most important checks on purpose, on a branch, and confirm each one turns the build red. A suite that catches every deliberate break is a real baseline. A suite that lets one or two through is telling you the actual starting point is weaker than the team believed, and that is exactly the kind of fact this whole exercise is built to surface early rather than after volume goes up.

Step three: the checks come first

This is the step that gets reversed, and reversing it is what turns a rollout into an incident.

The sequence that works is: get the blocking automated checks in place, confirm they can fail, then raise the volume of generated code. The sequence teams actually run is the opposite. Generation is immediately gratifying and check-writing is not, so the volume goes up first and verification is promised for later. By the time later arrives, there is a large body of code nobody can confirm is correct.

Before the volume goes up, the pipeline needs checks that block a merge rather than file a report, and each one has to have been seen to fail. Break a protected behaviour on purpose and confirm the build fails. Assertions that mock away the real path or compare a value to itself are common, and they pass forever. Adding evals to an existing codebase covers doing this on a codebase that starts with nothing.

The security case makes the ordering argument on its own: measured pass rates for AI-generated code are near 55 percent, and that is covered in is AI-generated code secure.

The two failure modes

The review queue becomes the slowest step. Generation gets faster, review does not, and within weeks review is the step that limits speed. The diagnosis is easy: throughput is flat while the number of open pull requests rises. The fix is not more reviewers. It is moving every question with a right answer into the automated suite, so reviewers spend their attention on design, scope, and whether the change was worth making. A reviewer reading a diff that already passes real checks is doing a different and much faster job.

The build passes and customers keep finding bugs. This one is more dangerous, because everything looks healthy. It means the checks are asserting the wrong things. The fix is to stop writing checks from the code and start writing them from real incidents: for each of the last ten things that reached a customer, write the check that would have caught it.

A worked version of that second failure mode makes it concrete. Say the incident was a discount code that applied twice when a customer changed their shipping address after checkout. A check written from the code would look at the discount function in isolation, confirm it multiplies correctly, and pass, because the function itself was never wrong. A check written from the incident starts from the sequence a real customer followed: apply a discount, change an address, resubmit, and assert the discount is applied exactly once at the end of that sequence. The first kind of check is comfortable to write and proves the least. The second kind is slower to write because it requires reconstructing what a user actually did, and it is the only kind that would have caught this specific bug before it shipped. A team that only ever writes the first kind can have hundreds of passing checks and still be exposed to the exact failure that just happened, because none of them asked the multi-step question.

A ninety-day plan

Weeks one and two, baseline the four metrics and pick the target service. Weeks three and four, write blocking checks for the highest-cost flows in that service and prove each one can fail. Weeks five to eight, run the new flow for real, spec first, generation second, checks deciding the merge, and let the team feel where it is awkward. Weeks nine to twelve, compare against the baseline and write down what you would do differently.

At the end you decide whether to extend it to a second team, on evidence. That decision is the entire point of the exercise, and it is the thing a rollout that buys licences and hopes cannot produce.

When not to do this

If you are mid-way through a launch with a fixed date, wait. If the codebase has no automated checks at all and no plan to add them, fix that first, because raising generation volume on an unverified codebase makes the problem worse at speed. And if the team is short-staffed enough that nobody can own the rollout, it will not happen. Adoption needs someone whose job it is, not a volunteer with a full sprint.

For what changes in the day-to-day once this is in place, what AI coding assistants actually change on an engineering team covers the human side, and what is AI-native software development sets out the definition the rest of this guide assumes.

Best for

  • An existing team with a working codebase, real users, and delivery commitments
  • Engineering leaders who need a defensible answer about whether AI tooling helped
  • Teams that already have some automated checks and can extend them

Avoid if

  • Do not start mid-launch with a fixed date: the deadline will take priority over the rollout
  • Do not start on a codebase with no automated checks and no plan to add them
  • Do not start without one named owner whose job the rollout actually is

Check before you decide

  • Confirm a delivery baseline exists before the first licence is issued
  • Confirm every blocking check has been seen to fail on purpose
  • Confirm the rollout is measured on delivery outcomes, never on volume of code or pull requests

Common questions

What should I measure before rolling out AI coding tools?

Take a two-week baseline of the four DORA keys: lead time for a change, deployment frequency, change failure rate, and time to restore. Do not measure lines of code, pull request counts, or commits, because those rise with AI adoption whether or not anything improved and will show a win in every scenario.

Why not just ask the team whether the tools are helping?

Because self-assessment is the one measure we know fails here. METR's July 2025 trial found developers took 19 percent longer with AI tools while estimating they had been 20 percent faster, a 39-point error in the direction that made the tools look better, among experienced engineers in code they knew well.

Should we roll out to the whole engineering organisation at once?

No. Change one flow in one repository with one owning team, and leave everything else alone. A change in a single place gives you something to compare against; the same change everywhere at once gives you impressions and no control group.

Do the automated checks really have to come before the volume?

Yes, and this is the step most often reversed. Generating code is immediately satisfying and writing checks is not, so verification gets promised for later, and by the time later arrives there is a large body of code nobody can confirm is correct. Get blocking checks in, prove they fail, then raise the volume.

What does it mean when throughput is flat but the number of open pull requests keeps growing?

The slowest step moved from writing to reviewing, which is the most common failure mode of a rollout. The fix is not hiring more reviewers, it is moving every question with a right answer into the automated suite so reviewers spend their time on design, scope, and whether the change was worth making.

Our build passes but customers still find bugs. What went wrong?

The checks are asserting the wrong things, which is more dangerous than having none because the dashboard looks healthy. Stop deriving checks from the code and derive them from real incidents instead: take the last ten problems that reached a customer and write the check that would have caught each one.

How long does an AI-native rollout take?

Plan about ninety days for one team: two weeks of baseline, two weeks writing blocking checks for the highest-cost flows, four weeks running the new flow for real, then four weeks comparing against the baseline. The output is a decision about whether to extend it, made on evidence.

When should a team not attempt this?

When a fixed launch date is close, when the codebase has no automated checks and no plan to add them, or when nobody can be given the rollout as an actual responsibility. The middle case matters most: raising generation volume on an unverified codebase makes the underlying problem worse, faster.