Guide

Building compliant software in financial services

Financial regulation is already written as rules. That is the part most teams miss. A rule that says a record must be recreatable after deletion is a specification, and a specification can be turned into a check that runs on every change. This guide covers how to do that, which obligations actually apply to your code, and why compliance gaps in a working app are so easy to miss.

Published August 22, 2026. Editorial.

Key takeaways

  • Compliance obligations are specifications. The ones that apply to your code can be written as executable checks, and a check that runs on every change is worth more than a policy nobody reads.
  • The SEC's audit-trail alternative under Rule 17a-4(f) requires recording who created, modified, or deleted a record, and being able to recreate the original. Most applications log the write and not the reader, and hard-delete on request.
  • The FTC Safeguards Rule names secure development practices, change management, and logging of authorized user activity as explicit requirements, which means the regulation covers your engineering process too.
  • AI-written code makes thorough checking affordable for the first time, which matters more in regulated work than anywhere else. It also produces insecure code by default, so the checks have to be independent of the code.
  • Most non-compliance in a working application is hard to see from inside the team. The app does what the team designed, and the gap sits in what nobody thought to design.

Ask a financial services team whether their application is compliant and you will usually get a confident yes, backed by a policy document, a completed questionnaire, and a penetration test from last year. Ask what in the running system enforces any of it and the answer is much less certain. That gap is the subject of this guide.

Here is the thing that makes regulated software different, and it is not difficulty. It is that somebody already wrote the specification. Most product work starts with a vague idea of what correct means, and the team discovers it by releasing the product. Regulated work starts with a rule, published in the Code of Federal Regulations, that states in operative language what the system must do. The SEC tells you an electronic recordkeeping system must maintain a time-stamped audit trail that includes the identity of the individual creating, modifying, or deleting a record. The FTC tells you that you must monitor and log the activity of authorized users. These are not values. They are requirements, and requirements can be tested.

That reframing is the whole argument for building regulated software the way we build it. When a rule is a specification, the right response is a check derived from that specification that runs on every change, forever. What almost every team does instead is read the rule once, implement something close to it, write a policy describing what they did, and then let three years of feature work slowly break it. The policy still says the right thing. The system stopped doing it eighteen months ago and nobody noticed, because nothing was checking.

Why this is the industry where AI-native development gives the most benefit

The economics of thoroughness changed, and regulated work is where that change is worth the most.

Writing four hundred compliance checks used to be a line item nobody approved. It was weeks of unglamorous engineering that delivered no features, so it lost every prioritisation meeting to the features customers could see. The result was that most financial applications ended up with a few tests for normal use and nothing at all for the obligations, which is the opposite of what a system whose main risk is regulatory needs.

When the code is close to free to write, that trade-off disappears. The four hundred checks are now an afternoon of specification work and a run, and they keep running on every change after that. Nothing else about our approach matters as much as this: the reason to build regulated software with AI is not that it is faster to release, it is that it makes exhaustive verification affordable, and exhaustive verification is what regulated work has always needed and rarely got. We wrote about the general form of this in adding eval tests was the best decision we made, and the full method is in our eval-driven development guide.

There is a second advantage that matters in examination. Regulators do not primarily want to know that your system is correct. They want evidence that you knew it was correct, at the time you released it, and can demonstrate that now. A suite of checks derived from the obligations, run automatically on every change, with timestamps and results retained, is that evidence as a byproduct of building. It is far more convincing than a policy document, because a policy describes intent and a passing check describes behaviour.

The honest limit

None of this works if you skip the part that makes it work, and we would rather say so plainly than hide it in a sales pitch.

Generated code is not secure by default. Veracode's spring 2026 testing across more than 150 models and 80 coding tasks found only 55 percent of generations produced secure code, a figure that has stayed roughly flat for two years while syntax correctness climbed above 95 percent [1]. Models learned to write code that compiles much faster than they learned to write code that is safe. In a regulated application, the paths where that matters most are exactly the ones the regulation cares about: authentication, authorisation, retention, and anything touching customer records.

So AI-native development without independent verification is worse than the traditional approach, not better, because it produces more of that code faster. The whole argument depends on the checks being derived from the specification before the implementation exists, and graded by something other than the thing that wrote the code. Our piece on why you should never let the model grade its own work covers the mechanism, and it is the difference between this being an advantage and being a liability.

What actually applies to your code

Financial regulation is enormous and most of it never affects an engineer's work. The useful step is to separate the obligations that are met by policy and process from the ones that can only be satisfied by the running system, and then treat the second set as engineering requirements with tests attached.

The obligations that apply to the code fall into a small number of areas. Recordkeeping and retention, covered on our SEC recordkeeping requirements in your application page. Access control and authentication. Audit logging, which is the single most commonly under-built area and has its own page on audit trails that satisfy an examiner. Encryption of customer information at rest and in transit. Change management, which is where the FTC Safeguards Rule applies directly to how you release software, covered on the FTC Safeguards Rule for product teams.

Each of those has a page in this guide that translates the rule into what the system has to do, and what a check for it looks like.

Why teams do not know their own app is non-compliant

The most useful thing we can tell you about this work is that non-compliance in a live application is usually hard to see from inside the team, and it is invisible for a structural reason rather than a careless one.

An application does what its team designed it to do. When the team reviews it, they review it against that design. The compliance gaps are almost never in the part somebody designed. They sit in the part nobody thought about: the reader who was never logged because logging was built for writes, the export path that bypassed the retention rule because retention was implemented in the delete button, the admin impersonation feature that acts as the customer and records the customer's identity in the audit trail rather than the admin's. Every one of those is a system working exactly as built. None of them was a decision.

This is why a self-review finds so little. You cannot review against a specification you never wrote down, and the whole reason these gaps exist is that the specification was never written down. Our page on the compliance gaps teams miss in their own app covers the specific patterns, and how to audit an existing financial application covers the process for finding them.

Which obligations apply to your business is a question for counsel rather than for engineers, and nothing in this guide is legal advice. What it does is take the obligations once somebody has scoped them and turn them into things a system can be tested against.

The frameworks this guide covers, and why those and not others

Financial services software sits under more than one regulator at once, and which ones apply depends on what the business does, not on what kind of company it calls itself. This guide covers the frameworks that reach an engineering team directly, because they name specific system behaviour rather than general risk management principles.

The SEC's recordkeeping rules, in 17 CFR 240.17a-4, apply to broker-dealers and set out what an electronic recordkeeping system must do to store and preserve records. The rule names the audit trail, the identity of the person acting, and the ability to recreate a record after it changes or is removed. Our page on SEC recordkeeping requirements in your application covers what that means for a data model and a storage layer.

The FTC Safeguards Rule, in 16 CFR 314.4, applies far more broadly than the SEC rule, because it reaches any business defined as a financial institution under the Gramm-Leach-Bliley Act, which includes many companies that do not think of themselves as financial institutions at all: a lender, a payment processor, a company that helps people file taxes, a business that arranges financing for its own customers. Section 314.4(c) is the part that names your development process specifically, and the FTC Safeguards Rule for product teams walks through what it requires.

SOC 2 is an attestation report, not a law, that a customer, usually an enterprise buyer or a bank the company is trying to integrate with, asks for before signing a contract. It is on this list because in practice it functions as a requirement: a growing financial software company will be asked for it whether or not any regulation demands it, and the underlying criteria overlap heavily with the SEC and FTC requirements above. What SOC 2 actually asks of your engineering process covers where that overlap is close enough to build once for both, and where it is not.

This guide does not cover the Bank Secrecy Act, anti-money-laundering program requirements, or state-level lending and licensing regimes, because those are largely policy and reporting obligations that sit above the application layer rather than inside it. A team that suspects those apply to its business needs counsel to scope them before an engineer can do anything useful with the answer.

How a compliant system stops being compliant without a single deliberate change

The failure pattern that matters most in this work is a system built right that stopped being right, one ordinary change at a time, with nobody deciding that it should.

Picture an application at the point it passes its first audit. The audit trail records every actor. Access follows the rule that a support agent can see only the accounts assigned to their queue. Every release goes through the same review and the same checklist. At that moment, the policy document and the running system agree with each other, and everyone involved has reason to believe the application is compliant, because it is.

Now run the ordinary life of a codebase forward for two years. A performance problem in the audit log leads an engineer to batch writes, and a batching bug quietly drops a percentage of entries during a deploy window, unnoticed because nothing was checking that every read produced a log line, only that the logging code existed. A new feature lets a support agent temporarily see an account outside their queue for the duration of an escalation, and the escalation path was built by a different engineer who did not know about the queue rule and had no reason to look for it, because it was never written down anywhere the second engineer would encounter it. A new release pipeline is adopted for speed, and the review step from the original checklist does not fit the new pipeline's shape, so it is quietly dropped rather than redesigned, because nobody owned the requirement that it exist. None of these three events was a decision to become non-compliant. Each was a normal, well-intentioned engineering choice made without visibility into a rule from a document nobody was looking at that day.

This is what makes compliance drift different from an ordinary regression. An ordinary regression breaks something a user notices, so somebody files a bug and it gets fixed. A compliance regression breaks a property that no user interaction depends on and no test exercises unless a test was written specifically to exercise it. The system keeps working. The feature the engineer shipped that day is fine. The only way to notice the drift is to keep re-testing the original requirement, on every change, indefinitely, which is exactly the work a policy document does not do and a check suite does.

The practical conclusion is that a point-in-time audit measures a system honestly and then goes stale immediately, at the same rate the codebase changes. The next section covers what replaces it.

Continuous verification, and what it actually replaces

A point-in-time audit answers one question: was this system compliant on the day someone with the right knowledge looked at it. That is a real answer and it is worth having, but it expires the moment the next change ships, because nothing about the audit process itself watches for the next change.

Continuous verification answers a different question: is this system compliant right now, and it answers it again automatically every time the system changes. The mechanism is the same one this guide has already described: an obligation is written as a specification, the specification becomes an assertion with an observable result, and the assertion runs in the same pipeline that runs before every release. The output is a pass or a fail attached to a specific commit, on a specific date, and a record of every time it ran, instead of a report filed once a year.

The difference this makes is concrete in the read-log example from earlier in this guide. A point-in-time audit, run once, would have found that the two or three code paths the original team built were logging reads correctly, and reported the system as compliant. It would not have found the internal support tool that shipped six months later and skipped the logging call, because the internal tool did not exist yet when the audit happened. A continuous check finds it on the day it ships, because the check tests the property, not the code path, and it runs against the new tool the same as it ran against the old one.

This does not mean a periodic external audit becomes worthless. An external reviewer brings a perspective the team building the system cannot fully have, because they are not the people who decided what to build and therefore do not share the team's blind spots. What changes is what sits between one external audit and the next. Under the old model, that gap was unmonitored, and a system could drift for a year and pass the next audit anyway if the drift happened to sit outside whatever the auditor chose to sample. Under continuous verification, the gap is covered by the same checks the whole time, so an external audit becomes a second opinion on top of a system that has been checking itself daily, rather than the only check the system gets.

How AI-generated code changes what needs checking

AI-assisted development changes the shape of the risk in a regulated codebase, and the change runs in both directions at once, which is why the two rules stated earlier in this guide, that the check comes from the obligation before the implementation and that the grading is independent of the code, are not optional extras.

In one direction, the risk goes up. A model given a task will produce a working answer to that task and nothing more, because it has no way to know about an obligation that was never stated in the prompt or the surrounding code it can see. Told to build a feature that exports account data to a spreadsheet, a model will build a feature that exports account data to a spreadsheet, correctly, and it will not independently notice that the export needs to honor the same retention and redaction rules as the screen the data came from, because nothing told it that rule exists. This is what happens when the only source of truth a system has is the task in front of it, across any model, and it is the same failure mode described earlier for a new engineer who was never told a rule existed, just faster and more frequent, because generated changes ship more often than human-authored ones did.

In the other direction, the same technology is what makes the fix affordable. Writing the check that catches the missed retention rule on the export path, before that feature is built, costs a small amount of specification work and a short run, for the reasons covered earlier in this guide. The same economics that make thorough verification affordable for a human-written codebase apply without change to an AI-written one, and they matter more there, because the volume of change is higher and the number of paths a specification-blind process can miss grows with it.

The practical result is that a regulated codebase built with AI assistance needs its check suite to arrive before the features do, not after. A team that writes the obligations as assertions first, then generates the implementation against a pipeline that already contains those assertions, gets a system where the model's blind spot is caught on the same run that introduces it. A team that generates first and writes checks afterward is testing the code against itself, which is the same circular problem described earlier in this guide with a test written after the implementation it is meant to verify.

What a real audit trail requires from the storage layer, not just the log line

Most teams that build an audit trail build a table that records events. Fewer build a table that survives the questions an examiner or an internal reviewer actually asks of it, and the gap between those two things is almost entirely in the storage layer rather than in the logging call.

An audit trail has to answer what a record looked like before the change, and that is the part a log line rarely covers. A log line that says a record was modified, without capturing the prior state, cannot satisfy the recreate-the-original-record language in the SEC's rule, because there is nothing left to recreate from. This means the audit trail needs its own storage, separate from the operational table it describes, holding either the full prior version of the record or enough of a difference from the current version to reconstruct it, and that storage has to be append-only in actual practice, enforced by permissions, rather than append-only only as a description in a policy document. A table that a sufficiently privileged account can update or delete rows from fails the requirement regardless of what the application code that normally writes to it intends, because the requirement concerns what the system permits, and permissions are the fact that governs it.

The trail also has to survive the deletion of the record it describes. A common design mistake links the audit entry to the record by a foreign key that cascades on delete, so deleting the customer record deletes the history that was supposed to prove what happened to it. The audit entry needs to be a durable statement that a record with a given identity existed and was acted on, independent of whether that record still exists, which usually means storing a copy of the identifying details in the trail entry itself rather than only a reference to a row that might later be gone.

Finally, the trail needs to record who, specifically, took the action, resolved to an individual, not a shared service account or an API key used by multiple people. A system that logs "API key 4" as the actor on a record deletion has recorded that an authenticated request happened, not who was responsible for it, and the identity requirement in the SEC rule and the monitoring requirement in the Safeguards Rule both specifically name the individual. Audit trails that satisfy an examiner covers the specific failure patterns in more depth, including the six ways a trail that looks complete on paper fails these three tests in practice.

If you are building or rebuilding a financial product, our fintech work and the broader custom software development service are the places to start. If what you actually need is to find out your current position first, get in touch and we will scope an audit against your real obligations rather than a generic checklist.

Compliance is a property of a running system, not a document you produce, and the only honest way to know you have it is to check.

Explore the guide

Common questions

Why is AI-native development a better fit for compliance-heavy software?

Because compliance obligations are already written as specifications, and the main cost of satisfying them has always been the unglamorous work of writing and maintaining exhaustive checks. When writing code is close to free, that work becomes affordable, and the check suite also serves as the evidence an examiner asks for. The advantage is thoroughness rather than speed.

Is AI-generated code safe to use in a regulated financial application?

Only with independent verification, and the evidence on this is clear. Veracode's spring 2026 testing found only 55 percent of model generations produced secure code while syntax correctness ran above 95 percent, so generated code compiles far more reliably than it protects. In regulated work the answer is to derive the checks from the obligation before the implementation exists and grade the result with something other than the model that wrote it.

What does the SEC actually require an application to record?

Under the audit-trail alternative in Rule 17a-4(f), an electronic recordkeeping system must maintain a complete time-stamped audit trail covering all modifications and deletions, the date and time of actions that create, modify, or delete a record, and the identity of the individual doing so, in a way that permits recreation of the original record. Most applications capture the change but not the actor, and hard-delete rather than preserving recreatability.

Does the FTC Safeguards Rule apply to how we build software?

Yes, and more directly than most teams expect. Section 314.4(c) explicitly requires adopting secure development practices for in-house developed applications that transmit, access, or store customer information, along with procedures for change management and controls to monitor and log the activity of authorized users. That places your development process itself inside what the regulation covers, not just the finished product.

How can an app be non-compliant without the team knowing?

Because the gaps sit in what nobody designed rather than in what somebody designed badly. Logging built for writes silently omits reads, retention implemented in a delete button is bypassed by an export path, and admin impersonation records the impersonated customer rather than the admin. Each is the system working exactly as built, which is why reviewing it against its own design finds nothing.

Is a penetration test enough to show our application is compliant?

No, because it answers a different question. A penetration test looks for exploitable weaknesses from the outside, while most compliance obligations concern what the system records, retains, and permits during entirely normal operation. An app can pass a penetration test cleanly while failing to log who read a customer record, which is a recordkeeping problem no attacker needed to be involved in.

Where should a team start if they suspect gaps but do not know where?

Start by writing down the obligations that apply to the running system, before looking at any code, so you have a specification to review against. Then test the four areas that account for most real gaps: who can read what, what is recorded when they do, what happens to data on deletion and export, and who approved the last ten changes to any of it.

Do these obligations apply to an internal tool nobody outside the company uses?

Frequently yes, and internal tools are where gaps concentrate because nobody treats them as products. If an internal tool transmits, accesses, or stores customer information, the Safeguards Rule requirements for access control, encryption, logging, and secure development apply to it, regardless of how few people use it or how quickly it was built.

Start a project