Building compliant software in financial services / Start here
Why AI-native development fits regulated work
The case for building regulated software with AI is not speed. It is that exhaustive verification finally became affordable, and regulated work is the place where exhaustive verification was always the actual requirement. This page explains the mechanism, and the condition the whole argument depends on.
Published August 22, 2026. Editorial.
Key takeaways
- Compliance obligations are published specifications, which makes them unusually well suited to being turned into executable checks.
- The historic reason regulated apps were under-verified is economic: the checks delivered no features and lost every prioritisation meeting.
- A check suite derived from the obligations also serves as examination evidence, because it shows behaviour rather than intent.
- This only works if the checks are independent of the code. Generated code is insecure by default, so self-graded work makes the problem worse.
Most arguments for AI in software development are about speed, and in regulated work speed is close to the least interesting benefit. A financial application that is released twice as fast and carries an unlogged read path is not a better outcome. The argument that matters here is different, and it is about what became affordable.
Regulated work has always needed something it could not pay for
Every serious financial application has the same kind of risk. The features are not especially hard. What is hard is that a long list of specific behaviours must hold, permanently, across every path through the system, including the paths added two years after anybody read the regulation. Who can see this record. What gets written down when they do. What happens when it is deleted. Whether the export respects the same rule as the screen.
The only reliable way to keep a list like that true is to check it automatically, on every change. Everybody in the industry knows this. Few teams did it, and the reason was never ignorance. It was that writing several hundred narrow checks over obligations is expensive, invisible, and delivers nothing a customer can see. Put that work in a prioritisation meeting against a feature the sales team is asking for and it loses, every quarter, until the compliance layer is a few checks for normal use and a policy document doing the rest of the work on paper.
That constraint is now gone. When implementation is close to free, the four hundred checks stop being a quarter of engineering budget and become a specification exercise plus a run. This is why we think regulated software is the single best fit for how we build: the thing AI made cheap is exactly the thing this industry always needed and could never justify.
Obligations are specifications, which is unusual and useful
Ordinary product work suffers from not knowing what correct means. The team has an idea, users have a different one, and the definition becomes clear over a year of releases. You cannot write a check against a requirement nobody has decided yet, which is why so much testing in normal software is limited.
Regulated work does not have that problem. Somebody already wrote the requirement, in operative language, and published it. The SEC's audit-trail alternative states that the system must maintain a time-stamped audit trail including the identity of the individual creating, modifying, or deleting a record, and must permit re-creation of the original record if it is modified or deleted [1]. That is not a value or a principle. It is close to a test case already, and turning it into one is routine work.
The same is true across the Safeguards Rule. "Implement policies, procedures, and controls designed to monitor and log the activity of authorized users" [2] is a sentence you can convert into a specific assertion about your system: for each role, for each customer record, a read produces a log entry naming the actor. Most teams read that sentence, build logging that covers writes, and move on. The sentence did not say writes.
This is the practical core of eval-driven development applied to a regulated context, and it is easier here than anywhere else, because the specification arrives written.
Evidence as a byproduct
There is a second benefit that only becomes obvious during an examination.
What a regulator wants is rarely a demonstration that the system is correct today. It is evidence that you knew it was correct when you released it, that you have known continuously since, and that you can show the working. Teams typically meet this with documents: a policy, a control matrix, a completed questionnaire, an architecture diagram. All of those describe intent.
A check suite derived from the obligations and run on every change produces something categorically stronger. It describes behaviour, with timestamps, over the whole history of the codebase. "Here is the assertion that a customer record read is logged with the acting user, here is every commit it ran against, here is the date it was introduced" is a different kind of answer than "our policy requires logging." It is also, usefully, an answer that cannot be produced retroactively, which is precisely why it is credible.
The condition the whole thing depends on
We would rather state the limit plainly than have you discover it.
Generated code is not secure by default. Veracode's spring 2026 testing across more than 150 models and 80 coding tasks found only 55 percent of generations produced secure code, and that figure has stayed roughly flat for about two years while syntax correctness climbed above 95 percent [3]. The gap is not narrowing. Models became excellent at producing code that runs and did not become good at producing code that is safe, and those were never the same skill.
In a regulated application, that matters most on exactly the paths the regulation cares about. Authentication, authorisation, retention, deletion, anything touching a customer record. So the naive version of AI-native development, where a model writes the feature and the tests in one pass and the pipeline passes, is worse than the old way. It produces more unverified code, faster, in the places where being wrong is most expensive.
Two rules make the difference, and neither is optional:
The check comes from the obligation, before the implementation exists. If you write the test after the code, the test encodes what the code does. That is circular, and in a compliance context it is worse than no test, because it creates false confidence with a passing result next to it.
The thing that grades the work is not the thing that did the work. A fresh context, ideally a different model, evaluating the diff against the written criteria. We covered the mechanism in never let the model grade its own work, and it matters more here than anywhere.
What does not change
Worth being clear about the limits. AI-native development does not decide which obligations apply to your business, and getting that wrong is a bigger risk than any implementation defect. It does not replace counsel. It does not decide whether a particular activity brings you into the scope of a particular regulation.
What it does is remove the excuse for the second half. Once somebody has decided what the rules are, there is no longer a good reason for those rules not to be enforced by something that runs on every change. The cost argument that justified the gap for twenty years is gone.
If you want to see how your current system compares against this, how to audit an existing financial application covers the process, and get in touch if you would rather we ran it with you.
Best for
- Teams whose main risk is regulatory rather than technical, where a long list of specific behaviours must hold permanently
- Products where the obligations are already written down and can be converted into checks
- Rebuilds and modernisations, where the check suite can be written from the obligations before the new implementation exists
Avoid if
- Nobody has decided which obligations apply, since a check suite cannot resolve a legal question
- You are not willing to fund the independent verification layer, in which case generating code faster makes the risk worse rather than better
Check before you decide
- Ask whether the checks were derived from the obligation or written after the implementation, and ask to see the order in the commit history
- Ask what grades the work, and whether it is the same context that produced it
- Ask which specific regulatory sentence each compliance check traces back to
Common questions
Is the argument for AI in regulated software mainly about releasing faster?
No, and speed is close to the least interesting benefit here. The argument is that exhaustive automated verification was always what regulated work needed and was too expensive to justify, and that constraint is now gone. A regulated app that is released faster with an unlogged read path is not a better outcome.
Why are compliance obligations easier to test than ordinary requirements?
Because somebody already wrote them down in operative language and published them, which almost never happens in ordinary product work. A rule stating that an audit trail must record the identity of the individual deleting a record is nearly a test case already, whereas a normal product requirement has to be discovered by releasing the product.
Does a passing check suite actually help during an examination?
It helps more than a policy document, because it demonstrates behaviour rather than intent, across the full history of the codebase with timestamps. It is also credible precisely because it cannot be produced retroactively, which a document can.
What is the single condition this approach depends on?
That the checks are independent of the code they verify: derived from the obligation before the implementation exists, and graded by something other than the thing that wrote it. Without that, generating code faster simply produces more unverified code in the paths where being wrong is most expensive.
How is this different from traditional manual code review for compliance?
Manual review depends on a reviewer remembering every obligation while reading a diff, which does not work for more than a few rules and rarely covers paths added later. A check suite derived from the obligations runs the same assertions on every single change, forever, so coverage does not depend on any one person's memory on any given day.
How would a team start applying this to an existing financial application?
Begin by listing the obligations that apply to the running system, quoting the regulatory text next to each one, before looking at any code. Then convert each into an assertion with a subject, an action, and an observable result, and write the check before touching the implementation, so the test can actually fail on first run.
Does building this level of verification take longer to release?
Writing several hundred checks from the obligations is a specification exercise measured in a short, defined amount of work rather than an open-ended one, and it runs automatically after that. The cost that used to make this impractical was the labour of writing and maintaining the checks by hand, and that is the part that changed.
Does this approach eliminate the risk of releasing insecure financial software?
No, and treating it that way would be the mistake. Generated code is not secure by default, so exhaustive checks reduce risk only when they are derived from the obligation before the implementation exists and graded by something other than the model that wrote the code. Skip that condition and the same volume of generated code simply carries more unverified risk, faster.
Related reading
Adding eval tests was the best decision we made
We generate every line of code we release. The single change that made that safe was writing the check before the code, and it turned out to change how we review, how we spec, and how fast we can work.
Never let the model grade its own work
When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.
How autonomous are AI coding agents, really?
Engineers at the leading AI labs now say a model writes one hundred percent of their code. Read the quotes closely and a person is still involved in every one of them. Here is what the 2026 evidence supports, and what it does not.