Who signs off on AI-written code?

The most useful question we get asked in a sales conversation is some version of this: if you generate all the code, who is actually responsible for it?
The answer is a person, by name, on every change. And the reason that matters is narrower and more important than it first sounds.
A check can only enforce a rule someone already thought of
Automated checks are the right answer to most of the quality problem. They are exact, they never get tired, and they keep up with code volume in a way human attention does not. We have written at length about building them, in the guide on eval-driven development, and we think a team releasing generated code without them is at risk in a way it probably has not measured.
But a check has a firm limit. It enforces a decision that was already made. It cannot make a new one.
So a suite will tell you that the permission rule you tested still works. It will not tell you that there is a second way into that data nobody thought about. It will confirm the export contains the columns you specified. It will not stop and say this report should not include salary data at all. It will prove the feature works. It will never say the feature was a bad idea.
Every one of those gaps is a judgment call, and judgment calls need an owner: a person who can be asked why.
What the named engineer is actually for
We have a specific view of what that person's job is, and it is not to re-read the code looking for bugs. If they are doing that, the automated checks are too weak and the real fix is earlier in the process.
Their job is four things.
Deciding this was the right change. Does it solve the problem the customer actually has, in a way that fits the product. This is the question that most often produces a rewrite, and no check can ask it.
Judging whether it will still be right in a year. Is this pattern one we want more of. Will this abstraction still make sense when the next three features are built on it. Generated code is uniformly tidy, which looks careful, so somebody has to ignore the tidiness and ask whether the structure is right.
Assessing whether the checks are strong enough. This is the part that surprises people. The most valuable thing a reviewer does on a change with good checks is look at the check itself and ask what it is not testing. A passing build against a weak test is the most dangerous state a change can be in, and we wrote about the specific ways that happens in when evals give false confidence.
Being answerable. The one that cannot be delegated to software. When a mistake reaches users, somebody gets asked why they thought this was fine. That question changes how people read a diff, weeks before it ever gets asked.
Why unclear ownership is the real risk
The failure we see is usually a decision that nobody owned.
The pattern is always similar. An agent wrote the change. Another agent, or the same one, wrote the tests. The pipeline passed. A reviewer approved it because the pipeline had passed and the code looked reasonable. Everyone involved behaved sensibly, and nobody at any point asked the question "is this actually correct, and how would we know". The change is released, breaks something that nobody notices at first, and the retrospective concludes that a test was missing. Which is true and useless, because the interesting question is whose job it was to notice that.
This is an old problem. Unclear responsibility has broken software for decades. What is new is how fast it now grows, and the fact that the volume of changes makes "everyone reviews everything carefully" arithmetically impossible.
Google's DORA research, across nearly 5,000 technology professionals in 2025, found AI adoption raising delivery throughput while lowering delivery stability, and described AI as something that makes an organisation's existing strengths and weaknesses larger. An organisation with clear ownership gets faster. An organisation where responsibility is shared so widely that nobody holds it gets faster at producing rework, and it takes longer to notice, because more is changing.
The uncomfortable part about self-review
There is one specific rule we hold and it is worth stating plainly: the thing that produced the work does not certify the work.
For humans this is old practice. Nobody approves their own pull request. For agents the same rule applies and it is easier to forget, because the reasoning that produced a change is still in the context that is now being asked to evaluate it. Anthropic's guidance for Claude Code recommends a fresh context or a separate subagent for exactly this reason, noting that a fresh context reviews better because it is not biased towards code it just wrote.
Our stronger version: the check is derived from the specification, before the implementation exists, so the standard cannot quietly change to fit the code. We wrote about the general principle in never let the model grade its own work.
What this means if you are buying software
Everyone uses AI now. A supplier who avoids it is slower rather than safer, so "do you use AI" tells you nothing. Four questions tell you a lot.
Who is the named engineer on our work, and will we meet them? Vague answers about a senior team are the ones to worry about.
What automated check must pass before a merge in our repository, and can we see it? Ask to see it, not to hear about it.
Are the checks written before or after the implementation, and by whom? Written after, by the same pass that wrote the code, means a suite that agrees with itself.
What happens internally when something reaches our customers? Listen for whether a person is involved or whether the answer is a process diagram.
That last question reveals the most. A supplier whose answer names a person has real accountability. A supplier whose answer names a process has only a procedure, and a procedure stops working under pressure. The wider version of this question set is in how to evaluate an AI development partner, and the business case for the automated checks is in the business case for evals.
Our answer, so you can check us against it
All of our code is generated. A named engineer reads every line before it reaches your branch, and that engineer is accountable for the judgment calls no check can make. The automated suite handles behaviour, contracts, unusual cases, permissions, and data invariants, so the human review is spent on whether this was the right thing to build and whether the checks around it are strong enough.
This does not make us perfect. It makes us answerable, which is the property you can actually verify before you sign anything.
Thanks to the clients who asked us this question directly and did not accept the first answer. The good version of our process exists because somebody challenged the vague version.
Software has always been accountable to a person. The only thing AI changed is how easy it became to lose track of which one.
Related guide: Eval-driven development: how to prove AI-written code works.
Sources
- Google Cloud, Announcing the 2025 DORA report: nearly 5,000 respondents; AI adoption shows a positive relationship with throughput and a negative relationship with delivery stability, with AI framed as an amplifier of existing organisational practice.
- Anthropic, Best practices for Claude Code: a fresh context or verification subagent so the agent doing the work is not the one grading it.


