Engineering

What a good technical spec looks like when a model writes the code

Editorial · Reveneau · August 2, 2026 · Updated August 19, 2026

What a good technical spec looks like when a model writes the code

For years the standard advice on technical specs was to keep them short. Name the problem, name the approach, name the risks, and stop where writing the code is the faster way to answer the question. That advice was right, and we gave it ourselves. It rested on an assumption nobody had to state, because it was always true: a person was going to read the document and build from it.

That assumption is the part that changed. When a model writes the implementation, the spec stops being only a briefing for a colleague and becomes the input the code is generated from. The old advice is still partly right, but it now divides into two rules that go in opposite directions.

How the cost of a gap changed

Start with what actually happens when a spec leaves a question unanswered, because that is the whole difference.

A human engineer reading your spec reaches the gap and notices it. Most of the time they ask. When they do not ask, they fill it with judgment: they know your business, they remember what happened last quarter, they have a sense that refunds probably should not be issued twice. Sometimes that judgment is wrong. But the decision is made by someone who can be asked about it afterwards, and who will often mention it in the pull request.

A model reading the same spec reaches the same gap and fills it with the most plausible answer. It does not stop. It does not say that a decision was made, because for the model there was no decision, only the next likely step. The gap is still there, hidden under something that looks reasonable. You find out what was chosen when the behaviour is wrong in front of a customer.

This is why leaving details out got more expensive rather than less. It used to be a reasonable trade: leave the small stuff open, let a good engineer handle it, save yourself a week of writing. Now the same choice quietly hands your product decisions to whatever pattern was most common in the training data. That is a bad place to make decisions about your business.

The split: behaviour got longer, implementation got shorter

The mistake is to conclude from this that specs should simply be longer. A spec contains two different things, and they moved in opposite directions.

Behaviour is what the system must do. What happens when the input is missing. What happens when the payment provider takes nine seconds to answer. What happens when the same request arrives twice because someone double-clicked. Which of two conflicting rules wins. This is the part to write out in far more detail than the old advice suggested, and it is the part most teams still skim, because it is tedious and it feels like paperwork.

This part is exactly where generated code fails. The first draft a model produces is usually correct about the main path and incomplete about everything around it, which is a specific and predictable weakness. Every unusual case you write down is a failure you do not release. Every one you leave out is left to chance.

Implementation is how the code is arranged. Variable names, file layout, function signatures, whether that helper lives in its own module. Specify less of this than ever. It was a waste of time when a person wrote the code, because a good engineer arranges code better than a document can, and it is still a waste of time now, for the same reason with a different worker. State the conventions if your codebase has them, then leave it alone.

The old rule was "stop where writing the code is faster than writing about it." The rule now is narrower and more useful: stop where the decision is about the structure of the code, and keep going as long as the decision is about what the software does.

What belongs in the document

Five sections still matter most, and one of them has grown.

The problem, and why it matters. Unchanged, still the most skipped, still the section that makes everything else understandable. If a reader does not understand the problem, they cannot tell a good solution from a plausible one.

The required behaviour. This is the section that grew. Write the main path, then write what happens at every point where it can fail or split into different cases. If you find yourself writing "handle errors appropriately", stop and write which errors and what appropriately means, because that sentence is precisely where a model will invent something.

The alternatives you rejected. Still essential, and worth more than it was. When the implementation can be regenerated cheaply, the temptation to try a different approach every time something gets hard goes up. A written record of what you already ruled out, and why, is what stops a team from arguing about the same decision again every sprint.

The risks and open questions. Name what you do not know. Admitted uncertainty is uncertainty you can plan around, which is the same argument we made from the buyer's side in how to write a software development RFP that gets honest bids.

The acceptance criteria. These matter more than they used to. They are what you check the generated code against, so write them concretely enough that they could become tests, because that is often exactly what happens to them.

A test for whether the spec is finished

Here is the useful part, and it is something the old process had no equivalent for.

Generate from your spec twice. Then compare what the two results actually do, not how they are written. Different variable names and a different file layout mean nothing. Different behaviour means something: it means your spec left that decision open, and something other than you made it, twice, in two different ways.

Every one of those differences is a sentence you need to write. The test takes an afternoon, it is repeatable, and unlike a spec review it does not depend on a reviewer happening to notice an absence. People are good at spotting a wrong statement and bad at spotting a missing one. This finds the missing ones.

Run it until the two generated versions agree on behaviour. At that point the spec is not perfect, but it is complete in the sense that matters: there is nothing important left for the model to guess.

Match the spec to the risk, still

None of this means every change needs a document. The size of the spec should still match the size of the risk, and for a small, well-understood change, writing one is overhead. That has not changed and it is not going to.

What changed is where the effort goes. The work that deserves a spec now deserves a more thorough one on the behaviour side, and a thinner one on the implementation side, and the total time is usually about the same or less. The hours moved to other work instead of increasing. They used to go into typing the implementation. Now they go into the first and last steps: deciding precisely what the software has to do, and checking that what came back does it. Generating the code in between costs almost nothing, which is exactly why quality now depends on the first and last steps.

A spec was always a tool for reducing uncertainty before it got expensive. It still is. The difference is that it is now read by something that will never tell you it was confused.

Sources

  • NIST, National Institute of Standards and Technology (research on the cost of finding and fixing software defects later in the software lifecycle): https://www.nist.gov/

Common questions

What is a technical spec?

A technical spec is a document written before building that explains what problem is being solved, how the team plans to solve it, what the system must do in each case, and what could go wrong. It has always helped a team think clearly and agree. When a model writes the implementation, it becomes something more direct as well: the input the code is generated from.

Has AI changed how long a technical spec should be?

It has changed which parts should be long. The sections describing required behaviour, including the unusual cases and the error cases, should now be far more complete than the old advice suggested. The sections describing implementation detail should be shorter than ever, because the model handles those better than you will.

Why does an ambiguous spec cost more when AI writes the code?

Because of how the gap gets filled. A human engineer who finds an unanswered question in a spec usually notices the gap and asks, or fills it using what they know about your business. A model fills the same gap with the most plausible answer and keeps going, without saying that a decision was made. The ambiguity stays, hidden until something behaves wrongly in production.

What should a modern technical spec include?

The problem and why it matters, the required behaviour in enough detail to cover the unusual and error cases, the alternatives considered and rejected, the risks and open questions, and acceptance criteria concrete enough to turn into tests. The acceptance criteria matter more than they used to, because they are what you check generated code against.

What should a technical spec still leave out?

Implementation detail: variable names, file layout, function signatures, and the internal structure of a module. Those were a waste of time to specify when a person wrote the code and they are still a waste of time now. Specify what the system must do, not how the code should be arranged.

How do you know a spec is detailed enough?

Generate from it twice and compare the behaviour of the two results. Anywhere the two differ in what the system actually does is a place your spec left a decision open and the model made it for you. That test is cheap, it is repeatable, and it shows the exact sentences that need work.

Should a spec include the error cases and unusual cases?

Yes, and this is the biggest change from older advice. Generated code is usually right about the main path and incomplete about everything else, so the parts of the spec covering what happens when an input is missing, a call times out, or two things happen at once are the parts that decide whether the output is usable.

Is writing a longer spec slower overall?

No, in our experience it moves the time to other steps rather than adding more. The hours that used to go into typing the implementation now go into the specification and the review, and the total is shorter because the step between them, where AI writes the code, costs almost nothing. A vague spec only delays the work until debugging.

Can a technical spec still change after it is written?

Yes, and it should. Building still teaches you things the document could not predict, and when what you learn while building disagrees with the spec, what you learned wins. The difference now is that the spec is a working document the code is generated from, so updating it is part of changing the system rather than paperwork you do afterwards.

Who should review a technical spec?

The people affected by the work and whoever owns the outcome, same as always. Add one more reader now: whoever will review the generated code. If they cannot tell from the spec what the correct behaviour is, they cannot judge whether what came back is right, and the review stage loses most of its value.