Building it right

When a human must stay in the loop

Most regulated systems have a human review step, and a significant share of those steps are not functioning as controls. A reviewer who approves nearly everything, at volume, without the information needed to disagree, is a formality that has a person's name on it. The difference is measurable, and worth measuring.

Published August 22, 2026. Editorial.

Key takeaways

  • A review that overturns the output in a negligible share of cases is not functioning as a review, whatever the process diagram says.
  • Measure your override rate. It is the clearest available signal of whether oversight is real.
  • Reviewers need the basis for the decision, not just the decision, and enough time per case to use it.
  • Design for the reviewer disagreeing: make override easy, recorded, and free of extra steps that discourage it.

Human oversight appears as a requirement or an expectation across regulated sectors, and it appears in most system designs as a step in a flow diagram. The gap between those two things is the subject of this page.

The formality problem

Consider a review step handling several hundred cases a day, where the model's output is presented as a recommendation, the reviewer has a queue and a target, and the interface shows the recommendation with an approve button.

That reviewer will approve nearly everything. Not because they are careless, but because the design gives them no basis for disagreement and no time to construct one. The recommendation arrives with the authority of the system, the volume makes deliberation impossible, and the easiest action is a single click.

The output of that arrangement is a system that makes decisions with a person's name attached. Whether that satisfies an oversight expectation is a legal question, and it should be an uncomfortable one, because the substance of oversight is not present.

A worked scenario shows how this happens without anybody deciding it should. A mortgage servicer builds a hardship review step where the model recommends whether a borrower qualifies for a payment plan, and a reviewer signs off before it goes into effect. In month one, the team staffs the queue generously and reviewers spend several minutes per case, checking the recommendation against the borrower's file. By month six, the volume has grown, the team has not, and the queue target has quietly become the metric that matters to the reviewer's manager. Nobody changed the policy. Nobody removed the review step. The number of seconds available per case simply fell until the only sustainable action was approval, and the process diagram from month one still describes the step accurately while the actual behaviour underneath it has become something else entirely.

Measure the override rate

The most useful diagnostic is also the simplest, and it is one almost nobody measures: what proportion of cases does the reviewer change?

If the answer is a small number, one of two things is true. The model is close to perfect, which is worth verifying independently rather than assuming. Or the review is not functioning, which is the more common explanation.

Measure it, segment it, and track it over time. Segment by reviewer, because a reviewer who never overrides while their colleagues sometimes do is telling you something. Segment by case type, because oversight may be real on the unusual cases and nominal on the routine ones, which is a defensible design if it is deliberate and a problem if it is accidental. Segment by time of day and queue depth, because override rates that fall sharply when the queue is long reveal what the review is actually sensitive to.

An override rate that is stable, non-trivial, and higher on hard cases is the typical sign of a working review. That is a measurable property, and it belongs in the check suite alongside everything else.

There is a failure mode on the other side worth naming, because a high override rate is not automatically good news either. If the rate is high and flat across every case type and every reviewer, that can mean the model's recommendations are simply not useful yet, in which case the review is working but the automation underneath it is not earning its cost. Distinguishing a healthy override rate from a symptom of a weak model requires looking at what the overrides actually change: a healthy pattern corrects specific errors on specific cases, while an unhealthy one looks like the reviewer redoing the work from nothing because the recommendation gave them nothing to build on.

What a reviewer needs

Three things, and most review interfaces provide the first only.

The basis, not just the output. What caused this result: the specific factors, the underlying data with its source and date, and what would have changed the outcome. A reviewer shown a score cannot evaluate it. A reviewer shown the two data items that produced the score can notice that one of them looks wrong.

Time proportionate to the consequence. If the decision materially affects someone, the review needs enough time to be a review. This is a staffing and throughput question that gets decided implicitly by queue design, and it deserves to be decided explicitly. A target that makes genuine review impossible has answered the question.

An easy way to disagree. If overriding requires a justification form, a supervisor approval, and a note, while approving requires one click, the design favours one outcome. Recording the reason for an override is right and necessary; making it burdensome is a way of discouraging the behaviour the control depends on.

Where oversight belongs

Not everywhere, and spreading it across every case is how it becomes nominal. The useful framing is to concentrate it where it does the most.

Adverse and irreversible outcomes. A decision that denies someone something, or that cannot be undone, is where review is worth its cost. Approvals and reversible actions can often proceed automatically.

Low-confidence cases. If the system can express uncertainty, route the uncertain cases to people. This is the highest-value routing available and it requires the model to have a calibrated notion of confidence, which is worth building for this reason alone.

Novel cases. Inputs unlike the training distribution are where models fail most and where a person adds the most.

Sampled routine cases. Even where routine cases proceed automatically, sampling some for review keeps the reviewers' judgment accurate and detects gradual changes. A review function that only sees hard cases loses its sense of what normal cases look like.

This concentration has a staffing consequence worth stating directly, because it changes what a review team looks like rather than just what it does. A team reviewing every case can be measured mainly on throughput, and hiring against a queue length is a familiar problem. A team reviewing only adverse, low-confidence, and novel cases is doing harder work per case by design, since the routine cases that would let a reviewer's attention rest have been routed away deliberately. Staffing and training the smaller team has to account for that: fewer reviewers handling harder cases need more time per case, not less, and a headcount plan built by dividing total volume by a fixed cases-per-hour rate will understaff a concentrated review function even though the total case count fell.

Design against automation bias

People defer to automated recommendations more than they should, and the effect is stronger when the system is usually right. A reviewer who has seen the model be correct a thousand times will not scrutinise the thousand and first carefully, which is precisely when the failure arrives.

Three design responses help. Present the evidence before the recommendation rather than after, so the reviewer forms a view first. Make the confidence visible, so uncertain cases feel different from certain ones rather than looking identical. And periodically include cases where the correct answer is known and differs from the model's output, which measures whether the review is functioning and keeps reviewers alert. That last one has to be handled carefully and transparently with the people involved, as a way to check accuracy rather than a trap.

A fourth response worth adding is rotating who reviews which case type over time, since a reviewer who has handled the same category for a long stretch develops exactly the confidence in the pattern that automation bias feeds on, only now the pattern is the reviewer's own prior judgment rather than the model's output. Fresh eyes on a familiar case type surface the same class of error that fresh eyes on a new hire's early decisions would, and the cost is only a rotation schedule rather than new tooling.

The honest position

Sometimes the right answer is that no human reviews each decision, and saying so is better than pretending.

A system that makes decisions automatically, with measurement, monitoring, an appeal path, and accountability for outcomes, may be a better design than one with a nominal reviewer who approves everything. It is certainly more honest, and honesty here has practical value, because a review step that exists on paper and not in substance is a weak position to defend and it also means nobody is actually checking.

What is not defensible is claiming meaningful oversight while the override rate says otherwise. If you are going to have a human review decisions, build the conditions that let them do the job, and measure whether they are doing it.

If you want help designing or measuring an oversight step, get in touch.

Best for

  • Systems where a model's output affects a person and a review step already exists on paper
  • Teams that want to know whether their oversight is real before somebody else asks

Avoid if

  • The decision is genuinely low-stakes and reversible, where automatic processing with monitoring is often the better design

Check before you decide

  • Measure the override rate overall, by reviewer, by case type, and under load
  • Check what the reviewer sees: the output alone, or the basis for it
  • Compare the effort required to approve against the effort required to override
  • Check whether low-confidence and novel cases are routed differently from routine ones

Common questions

How can you tell whether human oversight is real?

Measure the override rate, meaning how often the human reviewer changes what the system produced. A review that changes the output in a negligible share of cases is either checking a near-perfect model, which should be verified independently, or is not functioning as a review, and the second explanation is more common.

Why do reviewers approve nearly everything?

Usually because the design gives them no basis for disagreement and no time to build one. The recommendation arrives with the authority of the system, the volume makes deliberation impossible, and a single approve button is the easiest action, so the result is a system deciding with a person's name attached.

What does a reviewer actually need?

The basis rather than just the output, meaning the specific factors and underlying data with sources and dates; time proportionate to the consequence of the decision; and a path to disagree that is no more burdensome than agreeing. Requiring a justification form and supervisor sign-off to override while approval takes one click tells you which outcome the design prefers.

Where should human review be concentrated?

On adverse and irreversible outcomes, low-confidence cases where the system can express uncertainty, novel inputs unlike the training distribution, and a sample of routine cases to keep reviewers' judgment accurate and detect gradual changes. Spreading review across everything uniformly is how it becomes nominal.

Is it ever better to remove the human?

Sometimes, and saying so is more honest than pretending. A system deciding automatically with measurement, monitoring, an appeal path, and clear accountability can be a better design than one with a nominal reviewer approving everything, and it is a stronger position to defend than claiming oversight the override rate contradicts.

What is automation bias, and why does it matter for human oversight?

Automation bias is the tendency for people to defer to an automated recommendation more than the evidence warrants, and it grows stronger the more often the system has previously been right. A reviewer who has seen a model be correct many times in a row will not scrutinise the next case carefully, which is exactly the moment a failure is most likely to pass through unnoticed.

How do you measure an override rate in practice?

Record every case where a reviewer's final decision differs from the system's recommendation, then segment that count by reviewer, by case type, by time of day, and by queue depth. A rate that is stable and non-trivial, and higher on harder cases, is the typical sign of a review that is actually functioning rather than a formality with a person attached to it.

What does it cost to build a working oversight step versus a nominal one?

A nominal review costs almost nothing beyond an approve button and a queue. A working one costs staffing time proportionate to the consequence of each decision, an interface that shows the basis for a recommendation rather than just its output, and an override path with no more friction than approval, which together are a real ongoing cost rather than a one-time build.

Start a project