What a good technical spec looks like when a model writes the code

For years the standard advice on technical specs was to keep them short. Name the problem, name the approach, name the risks, and stop where writing the code is the faster way to answer the question. That advice was right, and we gave it ourselves. It rested on an assumption nobody had to state, because it was always true: a person was going to read the document and build from it.
That assumption is the part that changed. When a model writes the implementation, the spec stops being only a briefing for a colleague and becomes the input the code is generated from. The old advice is still partly right, but it now divides into two rules that go in opposite directions.
How the cost of a gap changed
Start with what actually happens when a spec leaves a question unanswered, because that is the whole difference.
A human engineer reading your spec reaches the gap and notices it. Most of the time they ask. When they do not ask, they fill it with judgment: they know your business, they remember what happened last quarter, they have a sense that refunds probably should not be issued twice. Sometimes that judgment is wrong. But the decision is made by someone who can be asked about it afterwards, and who will often mention it in the pull request.
A model reading the same spec reaches the same gap and fills it with the most plausible answer. It does not stop. It does not say that a decision was made, because for the model there was no decision, only the next likely step. The gap is still there, hidden under something that looks reasonable. You find out what was chosen when the behaviour is wrong in front of a customer.
This is why leaving details out got more expensive rather than less. It used to be a reasonable trade: leave the small stuff open, let a good engineer handle it, save yourself a week of writing. Now the same choice quietly hands your product decisions to whatever pattern was most common in the training data. That is a bad place to make decisions about your business.
The split: behaviour got longer, implementation got shorter
The mistake is to conclude from this that specs should simply be longer. A spec contains two different things, and they moved in opposite directions.
Behaviour is what the system must do. What happens when the input is missing. What happens when the payment provider takes nine seconds to answer. What happens when the same request arrives twice because someone double-clicked. Which of two conflicting rules wins. This is the part to write out in far more detail than the old advice suggested, and it is the part most teams still skim, because it is tedious and it feels like paperwork.
This part is exactly where generated code fails. The first draft a model produces is usually correct about the main path and incomplete about everything around it, which is a specific and predictable weakness. Every unusual case you write down is a failure you do not release. Every one you leave out is left to chance.
Implementation is how the code is arranged. Variable names, file layout, function signatures, whether that helper lives in its own module. Specify less of this than ever. It was a waste of time when a person wrote the code, because a good engineer arranges code better than a document can, and it is still a waste of time now, for the same reason with a different worker. State the conventions if your codebase has them, then leave it alone.
The old rule was "stop where writing the code is faster than writing about it." The rule now is narrower and more useful: stop where the decision is about the structure of the code, and keep going as long as the decision is about what the software does.
What belongs in the document
Five sections still matter most, and one of them has grown.
The problem, and why it matters. Unchanged, still the most skipped, still the section that makes everything else understandable. If a reader does not understand the problem, they cannot tell a good solution from a plausible one.
The required behaviour. This is the section that grew. Write the main path, then write what happens at every point where it can fail or split into different cases. If you find yourself writing "handle errors appropriately", stop and write which errors and what appropriately means, because that sentence is precisely where a model will invent something.
The alternatives you rejected. Still essential, and worth more than it was. When the implementation can be regenerated cheaply, the temptation to try a different approach every time something gets hard goes up. A written record of what you already ruled out, and why, is what stops a team from arguing about the same decision again every sprint.
The risks and open questions. Name what you do not know. Admitted uncertainty is uncertainty you can plan around, which is the same argument we made from the buyer's side in how to write a software development RFP that gets honest bids.
The acceptance criteria. These matter more than they used to. They are what you check the generated code against, so write them concretely enough that they could become tests, because that is often exactly what happens to them.
A test for whether the spec is finished
Here is the useful part, and it is something the old process had no equivalent for.
Generate from your spec twice. Then compare what the two results actually do, not how they are written. Different variable names and a different file layout mean nothing. Different behaviour means something: it means your spec left that decision open, and something other than you made it, twice, in two different ways.
Every one of those differences is a sentence you need to write. The test takes an afternoon, it is repeatable, and unlike a spec review it does not depend on a reviewer happening to notice an absence. People are good at spotting a wrong statement and bad at spotting a missing one. This finds the missing ones.
Run it until the two generated versions agree on behaviour. At that point the spec is not perfect, but it is complete in the sense that matters: there is nothing important left for the model to guess.
Match the spec to the risk, still
None of this means every change needs a document. The size of the spec should still match the size of the risk, and for a small, well-understood change, writing one is overhead. That has not changed and it is not going to.
What changed is where the effort goes. The work that deserves a spec now deserves a more thorough one on the behaviour side, and a thinner one on the implementation side, and the total time is usually about the same or less. The hours moved to other work instead of increasing. They used to go into typing the implementation. Now they go into the first and last steps: deciding precisely what the software has to do, and checking that what came back does it. Generating the code in between costs almost nothing, which is exactly why quality now depends on the first and last steps.
A spec was always a tool for reducing uncertainty before it got expensive. It still is. The difference is that it is now read by something that will never tell you it was confused.
Sources
- NIST, National Institute of Standards and Technology (research on the cost of finding and fixing software defects later in the software lifecycle): https://www.nist.gov/


