Building compliant software in financial services / Building it right
Using evals as compliance evidence
Compliance evidence is usually documentary: a policy, a control matrix, a completed questionnaire. All of it describes what you intended. A check suite derived from your obligations and run on every change describes what the system actually did, continuously, with dates attached. That is a different category of evidence.
Published August 22, 2026. Editorial.
Key takeaways
- Documents describe intent. Checks describe behaviour, which is what a control is supposed to produce.
- The evidential value comes from the history, and history cannot be manufactured retroactively.
- Every check should name the provision it satisfies, so coverage is visible and gaps are obvious.
- Checks that change need their own review trail, because weakening an assertion is the easiest way to make a failing pipeline pass.
There is a moment in most examinations where somebody asks how you know a control is working, and the answer is a document. The document says the control exists, describes what it does, and was last reviewed at some point. Everyone involved understands that the document is a description of intent, and that the actual question, whether the control was working continuously, is not answered by it.
A check suite answers it. That is the argument of this page.
Why behaviour is stronger evidence than intent
A control has two states in practice: what it was designed to do, and what it does. Documentary evidence captures the first. It is produced at design time, by the people who designed it, and nothing about the document changes when the system stops matching it.
That mismatch is the normal condition rather than the exceptional one. Systems change over time. Not through negligence, but because a feature added in month fourteen introduced a code path the original control did not anticipate, and nothing existed to notice. The policy still says customer record reads are logged. A new bulk export endpoint, added by a different team for a legitimate reason, reads through a path that predates the logging decorator. Both statements are true: the policy is accurate about intent, and the system no longer implements it.
An assertion that runs on every change closes that gap by design. The day the bulk export endpoint is added, the check that says every read produces a trail entry either covers it and fails, or does not cover it and shows its own missing coverage when somebody reviews it. Either outcome is better than silence.
The evidence is the history
The part that makes this genuinely strong evidence is not the check. It is the record of the check having run.
A passing test today tells you about today. Several thousand runs of the same assertion, attached to specific commits, across two years, tells you the control held over that period and lets you name the date it was introduced. That is close to the ideal answer to "how do you know this has been working," and it has a property that documents do not: it cannot be produced after the fact. A policy can be written the week before an examination. A two-year run history cannot.
This is worth designing for. Retain the results, not just the current state. Keep them for as long as the underlying obligation runs, which is usually far longer than a CI system's default retention. The number of teams who have excellent checks and a ninety-day results window is high, and the fix is a configuration change nobody thought to make.
Trace every check to a provision
A check suite is only usable as compliance evidence if somebody can map it to obligations. That means each compliance-relevant check carries a reference to the specific provision it exists to satisfy, in the code, next to the assertion.
The benefit is that coverage becomes a query rather than an exercise. Put the obligation list beside the check suite and the obligations with nothing attached are immediately visible. Without the trace, assessing your own coverage means inferring intent from test names, which is unreliable enough that most teams simply do not do it and assume they are covered.
It also makes amendments manageable. When a provision changes, one search finds the affected checks. This is not hypothetical: the proposed HIPAA Security Rule update would move encryption from an addressable specification to a required one, which changes what the assertion must say without changing what it is about. Traced, that is a targeted edit; untraced, it is a long search across a test suite nobody has read in a year.
Protect the checks themselves
Here is the failure mode that matters most, and it is not exotic.
When a pipeline fails and a release is waiting, there are two ways to make it pass. Fix the code, or weaken the assertion. The second is faster, and in a compliance suite it is catastrophic, because the assertion is the control. A test that has been weakened without notice still passes, still appears in the coverage map, and no longer verifies the thing it was created to verify.
This is not primarily about people acting in bad faith. An agent asked to get the build passing will take the quickest way to the stated goal, and so will a tired engineer at six on a Friday. The answer is structural rather than motivational:
Review changes to compliance checks separately from the change that motivated them, by someone who was not under the deadline. Fail the build when the number of skipped or disabled checks increases. Keep the assertions in files that require an additional approval to modify. Require a written reason, recorded, for removing or weakening any check traced to a provision. We wrote about the general version of this in never let the model grade its own work.
The point is that a compliance check suite is itself a control, and controls that can be silently disabled by the people they constrain are not controls.
What checks cannot do
Being honest about the boundary matters, because overselling this would be its own kind of compliance risk.
A check suite cannot tell you which obligations apply to your business. It cannot interpret an ambiguous provision. It cannot substitute for the judgment about whether a particular activity brings you inside the scope of a particular regulation, and it cannot cover an obligation nobody translated into an assertion. A suite is complete for what it was written to check and misses everything else, which means its coverage is exactly as good as the obligation list it was derived from.
So the sequence matters: counsel scopes the obligations, the team translates them into assertions using the method on how to translate a regulation into engineering requirements, and the suite enforces them from then on. Skip the first step and you get a rigorous, well-maintained, continuously verified implementation of the wrong requirements.
If you want help building the suite, or want to know what your current one actually covers, get in touch.
Common questions
Why is a check suite better compliance evidence than a policy document?
Because a document describes intent at the time it was written and does not change when the system stops matching it, while a check describes behaviour on every change. The gap between the two is the normal condition rather than the exception, since systems change as new code paths appear that the original control never anticipated.
What makes the evidence credible rather than just convenient?
The run history, which cannot be produced retroactively. A policy can be written the week before an examination, whereas thousands of runs of a specific assertion attached to specific commits over two years cannot be manufactured after the fact.
How long should we keep check results?
For as long as the underlying obligation runs, which is usually far longer than a CI system's default retention window. Plenty of teams have excellent checks and a ninety-day results window, which throws away the part that gives the evidence its value.
How do you stop someone weakening a compliance check so a release can go ahead?
Structurally rather than by instruction. Review changes to compliance checks separately from the change that motivated them, fail the build when the count of skipped checks rises, require an extra approval to modify the assertion files, and require a recorded written reason for removing any check traced to a provision.
What can a compliance check suite not do?
It cannot decide which obligations apply to your business, interpret an ambiguous provision, or cover an obligation nobody translated into an assertion. Its coverage is exactly as good as the obligation list it was derived from, so scoping has to come from counsel first.
Is an eval suite the same thing as a unit test suite?
Not quite. Ordinary unit tests mostly check that code behaves as the developer intended it to behave, and are often written after the code. An eval suite used as compliance evidence is written from the regulatory obligation before the implementation exists, traced to the specific provision it satisfies, and retained with its full run history rather than only its current pass or fail state.
How do you keep an eval suite from becoming a compliance exercise done only for show?
Guard changes to the checks themselves as carefully as the checks protect the code. Review modifications to compliance assertions separately from the change that motivated them, require a written reason to weaken or remove one, and fail the build when the count of skipped checks rises, so a shortcut under deadline pressure cannot silently disable a control.
Is it worth building an eval suite for compliance if the team is small?
The evidential value comes from the run history rather than team size, and a small team benefits the most, since it usually has the least spare capacity for a manual annual review. A modest suite that runs on every change and keeps its results starts accumulating credible history from day one.
Related reading
Adding eval tests was the best decision we made
We generate every line of code we release. The single change that made that safe was writing the check before the code, and it turned out to change how we review, how we spec, and how fast we can work.
Never let the model grade its own work
When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.
How many tests does AI-generated code need?
The honest answer is not a number, and coverage percentages are the wrong unit. Here is the unit we use instead, and why a suite of fifteen checks can be worth more than two thousand.
More in Building it right
Access control and least privilege in financial applications
The Safeguards Rule requires limiting authorized users' access only to the customer information they need to perform their duties. That is easy to state, hard to keep true, and almost never verified, because permission models get worse without anyone noticing and nothing breaks when they do.
Change management and release controls that keep working
Change management sounds like process overhead invented by people who do not release software. In regulated work it is a named requirement, and the version that satisfies a regulator is close to the version a good engineering team wants anyway: know what changed, know it was checked, know who agreed.
What SOC 2 actually asks of your engineering process
SOC 2 gets requested in nearly every enterprise financial deal, and treated by most engineering teams as a security questionnaire somebody else fills out. Read the actual criteria and it names your development process directly: change management, access control, and monitoring are not adjacent to SOC 2, they are most of it.