Build the eval suite

Writing specs an AI agent can verify

When a machine writes the implementation, the specification stops being a document people skim and becomes the actual input to the work. Vague specs used to produce slow projects. Now they produce large amounts of wrong code that looks right, fast. A spec that works has four parts: the behaviour stated as testable sentences, the boundaries named, the out-of-scope list written down, and an end-to-end check that proves the whole thing.

Published August 20, 2026. Updated September 30, 2026. Editorial.

Key takeaways

  • Ambiguity in a spec now turns into wrong code within minutes instead of being found in a status meeting weeks later.
  • Every requirement should be written so that its check is obvious. If you cannot see the check, the sentence is not finished.
  • Name the files, interfaces, and boundaries involved. An agent that has to guess the architecture will invent one.
  • State what is out of scope. Specs with no limits produce diffs with no limits.
  • End with one end-to-end verification that proves the feature works, rather than a list of components that exist.

The most common finding when we take over an AI-built codebase is code that faithfully implements a specification nobody had finished thinking about. The agent did what it was told. What it was told was incomplete, so it filled in the missing parts with plausible guesses and continued.

This part of the job got harder.

Why the spec matters more now

Two things changed at once. The implementation loop got fast, so a misunderstanding turns into thousands of lines before anyone reviews it. And the implementer stopped asking. A junior engineer with a vague ticket comes back with a question, which is a feedback mechanism nobody designed but everybody relied on. An agent does not come back. It picks an interpretation and continues.

So the ambiguity that used to be caught in conversation now has to be removed in writing. Anthropic's own guidance for larger features suggests having the model interview you first and then write the spec, and it names the property that matters: the most useful specs are self-contained, they name the files and interfaces involved, they state what is out of scope, and they end with an end-to-end verification step that proves the feature works [1]. Time spent making the spec precise is worth more than time spent watching the implementation.

That last sentence is the whole discipline.

Part one: behaviour as testable sentences

Write each requirement so that the check is obvious. The test is simple: can you tell, from the sentence alone, what input would prove it and what output would disprove it?

Weak: the export should handle large files properly.

Strong: an export of up to 100,000 rows completes inside the 30 second request timeout and returns a file whose row count matches the filtered query. Above 100,000 rows the request returns a job identifier instead, and the file arrives by email within ten minutes.

The second version is longer, and every extra word is needed. It contains four decisions that somebody was going to have to make anyway, and it makes them at the cheapest possible moment. It also produces its own evals directly, which is the point: the sentences in the spec become the checks in the suite, one to one where possible. The mechanics of turning them into checks are in how to write your first eval suite.

Part two: name the boundaries

An agent given a task and no architectural context will invent structure. Usually reasonable structure, and usually not yours. Then you have two patterns in the codebase doing the same job, which is a maintenance cost that grows over time without anyone noticing.

So name things. Which module owns this. Which existing function it should extend rather than duplicate. Which interface it must satisfy. Which pattern already in the repository is the example to follow. Where the tests live. Anthropic's guidance on prompting says the same thing in practice: point to the existing pattern in the codebase and name a good example file rather than describing the outcome and hoping [1].

This is also where a project-level convention file is useful, because it stops you rewriting the same context in every task.

Part three: write down what is out of scope

Specs with no limits produce diffs with no limits. An agent asked to add a feature will often improve nearby code while it works, and a 200-line change becomes an 800-line change that touches three modules and needs a much more careful review.

So state it. Do not change the authentication flow. Do not upgrade the dependency. Do not reformat files you are not editing. Do not add a caching layer. If a refactor turns out to be necessary, raise it and stop.

The out-of-scope list is the cheapest review-time saver we know. It turns a kind of surprise you would otherwise find in a diff into a question the spec has already answered.

Part four: one end-to-end verification

Finish with the single check that proves the feature works as a whole, expressed as something runnable rather than as a description. For example, "running the export command against the seed dataset produces a file with 4,312 rows and exits zero", rather than "the export feature is complete".

This matters because of how an agent decides it is finished. It stops when the work looks done, and without a check it can run, looks-done is the only signal available [2]. A named end-to-end verification gives the run a real point at which to stop, and it gives your reviewer a one-line answer to the question of whether this actually works.

What good looks like in practice

A working spec for a normal feature is usually one page. Five to ten behaviour sentences that each imply a check. A short list of files and interfaces. Three or four out-of-scope lines. One end-to-end verification. Written in twenty minutes to an hour, and it removes most of the rework that a vague version would have produced.

It also has a side effect worth naming. Because the sentences are testable, the spec and the eval suite stop being separate artefacts. The suite becomes the executable version of the spec, which means it stays current, which means the documentation problem partly solves itself. That is rare enough to be worth the discipline on its own.

Where this connects

The counterpart to a precise spec is a check that enforces it, so read what to check in an eval suite next. If the real problem is that nobody knows what the product should do at all, that is an earlier problem and knowing what to build and what a good technical spec looks like are the better starting points.

Common questions

How do I write a spec an AI coding agent can act on?

Write each requirement as a sentence whose check is obvious, name the files, interfaces, and existing patterns involved, list what is out of scope, and end with one runnable end-to-end verification. A page is usually enough for a normal feature, and it removes most of the rework a vague version would cause.

Why does spec quality matter more with AI-written code?

Because the feedback mechanism disappeared. A person with a vague ticket comes back to ask a question, while an agent picks a plausible interpretation and implements it at speed, so ambiguity becomes thousands of lines of wrong code that looks right before anyone reviews it.

Should the spec include the tests?

It should include the behaviour sentences that the tests are derived from, ideally one to one. When the spec is written that way the eval suite becomes the executable version of the spec, which keeps both current instead of letting the document fall out of date.

Why write down what is out of scope?

Because an agent asked to add a feature will often tidy nearby code while it works, turning a small change into a large one that needs a much longer review. A few out-of-scope lines turn that surprise into a decision the spec already made.

How long should a spec for an AI coding agent be?

A working spec for a normal feature is usually one page: five to ten behaviour sentences that each imply a check, a short list of files and interfaces, three or four out-of-scope lines, and one end-to-end verification. It typically takes twenty minutes to an hour to write and removes most of the rework a vague version would cause.

What is an end-to-end verification in a spec?

It is one runnable check that proves the whole feature works, stated as something concrete rather than a description, such as running a specific command against a seed dataset and confirming the output matches an exact expectation. It matters because an agent stops when the work looks done, and a named verification gives it a real point at which to stop instead of a guess.

Why should a spec name specific files and interfaces?

Because an agent given a task with no architectural context will invent a reasonable-sounding structure that is not necessarily yours, producing two patterns in the codebase doing the same job. Naming which module owns the work, which function to extend, and which existing pattern to follow removes that guesswork before it becomes a maintenance cost.

Does writing a detailed spec replace the eval suite?

No. The spec is what the eval suite is built from. A spec written as testable sentences produces the eval suite almost directly, since each sentence implies a check that can be written one to one against it. The spec and the suite become the same discipline seen from two sides: one states the requirement, the other proves it holds.