Eval-driven development: how to prove AI-written code works / Run it continuously
Running evals in CI when agents write the code
An eval that does not run on every change is only a note somebody wrote once. The value comes from blocking: no change reaches the main branch unless the relevant checks passed. That is simple to say and has a few real design decisions inside it, mostly about speed, about what an agent is allowed to do with a failed build, and about what happens on the day the check is wrong.
Published August 20, 2026. Updated September 30, 2026. Editorial.
Key takeaways
- Blocking bad changes is the whole purpose. A suite that cannot block a merge will not be maintained.
- Split the pipeline by speed so the fast checks give feedback in minutes and the slow ones still block the merge.
- Give the agent the same command your pipeline runs, so it fails locally before it fails in CI.
- Never let an agent disable, weaken, or skip a check to get a passing result. That rule needs to be enforced by the system.
- Every defect that reaches users should become a new check in the suite within the week.
Continuous integration is not new and most teams have some. What changes with generated code is the volume passing through the checks and the fact that one of the contributors is a machine that will optimise for whatever signal you give it.
Make the block real
Start here because everything else is detail. A merge into your main branch should be impossible while the relevant checks are failing, enforced by branch protection settings rather than by team habits.
Teams resist this because of the day the check is wrong and blocks urgent work. That day comes. Handle it with a documented override that requires a second person and leaves a record. Making the block optional is the wrong answer: an optional block weakens without anyone noticing, because every individual decision to merge past it is defensible and the total result is a suite nobody believes.
Split by speed
The practical constraint is that people find ways to avoid a slow pipeline. Our default split is three tiers.
On every push, in under two minutes: type checking, linting, and the unit-level evals. This is the tier the agent uses as its feedback loop, so it has to be fast.
Before merge, in under fifteen minutes: integration evals, contract and schema checks, permission and tenancy cases, migration checks against data with the same structure as production. These are the ones that catch the expensive mistakes.
On a schedule, and after deploy: the long ones. Full end-to-end journeys, performance thresholds, security scans, and any model-graded checks that cost real money to run. Nightly is fine for these, as long as a failure creates work rather than an email nobody opens.
The tiering rule is about waiting time and cost. A permission check is more important than a lint rule and it belongs in the slower tier because it takes longer to run.
Give the agent the same command
The pipeline and the local run should invoke the same thing. One command, checked into the repository, that runs the tier you want. If the agent's local check differs from the pipeline's check, it will produce work that passes locally and fails in CI, and every one of those round trips wastes a human's attention.
This is the mechanism behind the advice to give an agent a way to verify its work: the check has to be something it can run and read [1]. Two stronger options exist once the command exists. You can make the check a stop condition, so the agent keeps working until it passes rather than stopping when the work looks done. And you can enforce it as a hook that blocks the turn from ending, which is the deterministic version [1].
The effect on a working day is large. The agent finds its own failures, tries again, and gives you something that passes. You spend your attention on whether it was the right change.
The rule about not weakening checks
An agent optimises for the signal you give it. If the signal is a passing pipeline, then deleting an assertion, adding a skip, loosening a threshold, and widening a mock are all valid strategies. They are simply the shortest way to reach the goal you stated.
So the goal has to be stated properly, and it has to be enforced somewhere the agent cannot reach:
Changes to the eval suite are reviewed separately from the change that motivated them. A diff that touches both an implementation and the tests protecting it gets read with that in mind.
Deleting or skipping a check requires a written reason. In practice a pipeline step that fails when the count of skipped tests rises is worth more than a policy document.
Thresholds live in a file that is protected, not in the test that uses them.
Coverage minimums, if you use them at all, can only go up.
This is the same reason we do not let anyone approve their own pull request.
What to do when a check fails
Three outcomes, and naming them makes the team calmer.
The code is wrong. Fix the code. This is the common case and the system worked.
The check is wrong. Fix the check, in its own change, with a reason. It happens, particularly with new checks and time-dependent data.
The check is unreliable. This is the dangerous one. Do not re-run it. Take it out of the blocking tier the same day, then fix it or delete it. A flaky check left in place turns your whole suite into advice that nobody has to follow. That failure mode has its own page: when evals give false confidence.
Feeding the suite from production
The last loop is the one that keeps the suite accurate. When something reaches a user, the eval that would have caught it gets written before the fix, and it stays forever.
This turns your incident history into an automated check that never forgets, and it means your suite ends up matching your product's real failure modes rather than a testing textbook. Over a year that is the difference between a suite that catches things and a suite that exists.
For what to measure once the pipeline is doing this work, see metrics for AI code quality.
A worked example of the tiering decision
Say a team is adding a permission check to an existing endpoint: a user should only be able to cancel their own subscription, not anyone else's. Three checks could catch a regression here, and they belong in three different tiers for reasons that have nothing to do with how important each one is.
A type check that the endpoint's handler receives a user identifier of the expected shape runs in under a second and belongs on every push, because it costs nothing to run constantly and catches an entire category of mistake before anything else runs.
An integration check that calls the real endpoint as a different user and confirms the cancellation is refused takes a few seconds against a real database and belongs in the pre-merge tier, because it needs infrastructure that is not worth spinning up on every keystroke but is cheap enough to run before every merge.
A full end-to-end check that walks through account creation, subscription, and cancellation as two separate browser sessions, confirming the whole journey end to end, might take two minutes and belongs on a schedule, because it is the most convincing proof and the most expensive one, and running it on every push would make the fast loop unusable for everyone working on unrelated code that day.
The permission requirement is the same in all three. What changes is how expensive proving it is, and the tiering rule sorts by that cost rather than by which check feels most important. A team that skips the tiering and puts everything in the pre-merge tier ends up with a fifteen-minute wait for every change, and people start finding ways around it. A team that puts everything in the fast tier ends up with tests that cannot afford to touch real infrastructure, so the permission check either does not exist or is faked with a mock that proves nothing.
What happens when an agent meets a slow or unclear pipeline
An agent given a vague or slow feedback loop behaves in a specific, predictable way: it optimises for whatever makes the loop go green fastest, and a slow pipeline pushes it toward doing that with less information. If the fast tier does not actually exercise the code the agent just wrote, the agent will see a green result after making a change that the fast tier was never able to catch, try a much larger change on the strength of that false confidence, and only discover the real problem fifteen minutes later when the pre-merge tier finally runs. Every one of those fifteen-minute round trips is an iteration the agent could have used to fix the actual problem instead.
This is the practical reason the fast tier needs to be genuinely representative of common failures, not just fast. A fast tier that only runs a linter gives quick feedback about nothing that matters. The right fast tier includes the unit-level evals for the code most likely to be touched that day, chosen so that a wrong turn gets caught in the same loop the agent is already running rather than fifteen minutes and one context switch later.
Handling the case where the agent cannot fix the failure
Sometimes a check keeps failing and the agent's attempts do not converge on a fix, either because the requirement was stated ambiguously or because the fix genuinely needs a decision only a person can make, such as choosing between two valid designs. The correct behaviour here is for the agent to stop and describe what it tried and why each attempt failed, rather than to keep guessing or, worse, to modify the check itself to make the guess pass. This is the same rule as the one about not weakening checks, applied to the moment when it is most tempting: an agent under a Stop hook that cannot end the turn until the check passes has a real incentive to find the weakest possible way to satisfy that condition, and the review-the-suite-separately rule exists specifically to catch that.
Ownership of the pipeline configuration itself
One detail that is easy to miss: the CI configuration, the branch protection settings, and the list of required checks are themselves part of the codebase, and they deserve the same scrutiny as the code they gate. A change that quietly removes a check from the required list is functionally identical to a change that deletes the check outright, and it should be reviewed with the same suspicion. Some teams put pipeline configuration behind a separate approval path from ordinary code changes for exactly this reason, so that loosening what blocks a merge is never a side effect of an unrelated pull request.
Best for
- Any repository where an agent contributes code, however closely supervised
- Teams whose review queue is the slowest step, rather than implementation
- Codebases with expensive failure modes: payments, permissions, customer data
Avoid if
- Do not run a suite advisory-only: a check that does not block soon gets ignored
- Do not put slow checks in the fast tier and make the loop unusable
- Do not let the same change weaken a check and claim the passing build it produced
Check before you decide
- Confirm branch protection actually prevents a merge with failing required checks
- Confirm the local command and the pipeline command are the same thing
- Confirm skipped or deleted checks are visible and require a written reason
Common questions
How should evals be connected to CI?
Make the relevant checks required for a merge, enforced by branch protection rather than by convention, and split the suite into tiers by speed: fast unit-level checks on every push, integration and permission checks before merge, and long end-to-end, performance, and model-graded checks on a schedule. Keep the local command identical to the pipeline command.
What stops an AI agent from deleting a failing test?
Enforcement rather than instruction. Review changes to the suite separately from the change that motivated them, fail the pipeline when the number of skipped checks rises, keep thresholds in protected files, and require a written reason for removing a check.
What should happen when a check fails intermittently?
Remove it from the blocking tier the same day, then fix it or delete it. Re-running a flaky check teaches the whole team that a failure does not mean anything, which turns every other check into advice without anyone noticing.
How fast does the pipeline need to be?
The tier an agent or engineer uses as a feedback loop should return in about two minutes, and the pre-merge tier within roughly fifteen. Beyond that people start avoiding it, and a check people avoid blocks nothing.
Should the same command run locally and in CI?
Yes, one command checked into the repository should run the tier being tested, both locally and in the pipeline. If an agent's local check differs from the pipeline's check, work will pass locally and fail in CI, and every one of those round trips wastes a human's attention on a mismatch that should not exist.
What should happen when a required check turns out to be wrong?
Fix the check in its own separate change, with a written reason attached, rather than loosening it without notice inside the change that triggered the failure. This happens most often with new checks and time-dependent data, and keeping the fix isolated preserves a record of why the check changed.
How should an urgent merge be handled when a check blocks it?
With a documented override that requires a second person and leaves a record, rather than by making the check optional. An optional check weakens without anyone noticing because each individual decision to merge past it looks defensible on its own, and the total result becomes a suite nobody actually trusts.
Why split the pipeline into speed tiers instead of by check type?
Because the practical limit is waiting time: people avoid a slow pipeline regardless of how important its checks are. Fast unit-level checks run on every push, integration and permission checks run before merge, and long end-to-end or security checks run on a schedule, sorted by how long they take rather than by how important they are.
Related reading
How to move fast without breaking the product
Speed and quality are usually described as a trade-off. In practice, the teams that work fastest over a long time are the ones that made quality cheap to keep.
How to onboard engineers fast so they release work in week one
Slow onboarding is a hidden cost you pay on every hire. Here is how to get a new engineer releasing real work in their first week instead of their first month.
More in Run it continuously
Metrics for AI code quality: what to watch instead of volume
Once code became cheap to produce, every metric based on how much of it you produce stopped telling you anything useful. What still means something is what comes back: how often a change breaks something, how long recovery takes, how many defects reach a customer, and how much of last month's work is being redone. Those four are still useful after the change, and the first two have ten years of research supporting them.
When evals give false confidence
A suite that catches nothing is worse than having no suite, because a team with no checks knows it is exposed and a team with passing checks believes it is covered. Six failure modes account for almost all of it, and each one has a specific sign you can look for this afternoon.
Agent debt: when generation outpaces review
Teams now generate most of their code with an agent and review almost none of it as carefully as before, because the agent is fast enough to make thorough review feel like the slowest step. The cost of that missing review appears later, in production, months after the pull request was merged, as an incident nobody can trace to a specific change.