Guide

Software for compliance-heavy industries

Regulated software has always needed exhaustive verification and rarely received it, because writing the checks cost more than anyone would approve. That constraint is gone. This guide covers what that changes, the obligations that apply regardless of sector, and how to tell a partner who verifies their work from one who says they do.

Published August 22, 2026. Editorial.

Key takeaways

  • Compliance obligations are published specifications, which makes them unusually suited to being turned into checks that run on every change.
  • The historic reason regulated systems were under-verified was economic, and that reason no longer holds.
  • Generated code is insecure by default, so AI-native development without independent verification is worse than the traditional approach rather than better.
  • A check suite derived from obligations is stronger evidence than a policy, because it demonstrates behaviour over time rather than intent at a point.
  • When choosing a partner for regulated work, ask what grades the code and in what order the checks were written. The answers separate practice from marketing.

There is a claim we make about our own work that sounds like marketing until you look at where the cost actually is: regulated industries are the place where AI-native development gives the most benefit, and the reason has nothing to do with releasing faster.

Every serious regulated system has the same structure. The features are usually not difficult. What is difficult is that a long list of specific behaviours must hold, permanently, across every path through the system, including the paths added two years after anybody last read the regulation. Who can see this record. What gets written down when they do. What happens when it is deleted. Whether the export obeys the same rule as the screen.

The only reliable way to keep a list like that true is to check it automatically on every change. Everybody in these industries knows it. Few teams did it, and not out of ignorance. Writing several hundred narrow checks over obligations is expensive, invisible to customers, and produces nothing anybody can demo. Compare that with a feature the sales team is asking for and the checks lose, every quarter, until the compliance checks cover only normal use and a policy document covers the rest on paper.

That constraint no longer exists. When implementation is close to free, the four hundred checks stop being a quarter of an engineering budget and become a specification exercise plus a run. The thing AI made cheap is precisely the thing regulated work always needed and could never justify.

Compliance obligations are specifications, and that is unusual

Ordinary product work suffers from not knowing what correct means. The team has an idea, users have another, and the real definition emerges over a year of releases. You cannot write a meaningful check against a requirement nobody has decided, which is why testing in most software is shallower than anyone would like.

Regulated work does not have that problem. Somebody already wrote the requirement, in operative language, and published it. When the SEC says an electronic recordkeeping system must maintain a time-stamped audit trail including the identity of the individual creating, modifying, or deleting a record, and must permit re-creation of the original [1], that is close to a test case already. When the FTC says you must monitor and log the activity of authorized users [2], that converts into a specific assertion about your system: for each role, for each customer record, a read produces an entry naming the actor.

Most teams read that second sentence, build logging that covers writes, and move on. The sentence did not say writes. The gap between those two things is where a great deal of real non-compliance is found, and it is the kind of gap a check finds and a review does not.

The limit, stated plainly

None of this works if you skip the part that makes it work, and we would rather say so than hide it in the sales pitch.

Generated code is not secure by default. Veracode's spring 2026 testing across more than 150 models and 80 coding tasks found only 55 percent of generations produced secure code, a figure roughly flat for two years while syntax correctness rose above 95 percent [3]. Models became excellent at producing code that runs and did not become good at producing code that is safe. Those were never the same skill, and the gap is not closing.

In a regulated system that failure rate affects the paths the regulation cares about most: authentication, authorisation, retention, deletion, anything touching a customer or patient record. So the naive version of AI-native development, where a model writes the feature and the tests in one run and the pipeline passes, is worse than the old way. It produces more unverified code, faster, exactly where being wrong is most expensive.

Two rules make the difference and neither is optional. The check comes from the obligation, before the implementation exists, so the standard cannot be copied from the thing being measured. And the thing that grades the work is not the thing that did the work, which we cover in never let the model grade its own work.

What this guide covers, and what the industry guides cover

This hub contains the argument and the obligations that reach every regulated product regardless of sector: the EU AI Act, enterprise security review, retention and deletion across copies, and what an audit-ready codebase looks like. It also covers how to choose a partner for this work, which is a harder question than it looks because every supplier now says they test.

The sector-specific material is in three companion guides. Recordkeeping, audit trails, and the FTC Safeguards Rule are in financial services compliance. PHI, the Required and Addressable distinction, and minimum necessary are in healthcare compliance. Algorithmic fairness, valuation model standards, and screening are in real estate compliance. This hub carries the parts that cross all of them: how to turn any obligation into a check at all, on compliance is a specification problem; the EU AI Act's deferred timeline, on the EU AI Act and your product; retention and deletion across every copy of your data, on retention and deletion across every copy; the structural properties that make a codebase auditable, on building an audit-ready codebase; when a human review step is real rather than nominal, on when a human must stay in the loop; and how to vet a partner for this specific kind of work, on choosing a development partner for regulated work.

What the naive approach looks like

Nearly every regulated company we talk to has already built something that looks like compliance infrastructure, and it is worth describing precisely what that infrastructure usually is, because the description is the same across sectors even though the regulations differ.

The naive approach treats compliance as a documentation exercise that runs alongside engineering rather than through it. A compliance team, sometimes one person wearing several hats, reads the applicable regulations and writes policies describing how the organisation will meet them. Engineering builds the product according to feature requirements that rarely reference the policy document directly. Periodically, usually annually or before a certification renewal, someone maps the policy against the system, finds it is mostly right, patches the visible gaps, and files the updated policy. Everyone involved has behaved reasonably. The organisation can produce a document for anyone who asks.

This feels sufficient because it produces the artefact regulators, auditors, and enterprise customers historically asked for: a policy. For a long time, a policy was treated as adequate evidence because verifying the running system against the policy in detail was too costly for anyone to do, so nobody, including the regulators, expected it.

Why it fails in practice

It fails because a policy describes intent at the moment it was written, and a codebase changes underneath it continuously without producing any signal that the two have diverged. The failure has a specific mechanical shape, and it is the same shape in every regulated industry we cover on this site.

An obligation is written in general language: "monitor and log the activity of authorized users" [2], "maintain a complete audit trail" [1], "comply with applicable nondiscrimination laws." Someone translates that into a policy sentence, and the translation quietly narrows the scope, almost always without anyone intending it to. "Monitor activity" becomes "log changes," because logging changes is what the team already knew how to build, and logging reads as well felt like a separate, harder project that got deprioritised. The policy document, read on its own, sounds like it satisfies the regulation. The running system, examined directly, does something narrower. Nobody notices the gap because nothing forces a comparison between what the policy says and what the code actually does, and the two documents, the regulation and the codebase, never talk to each other.

This is the same failure whether the obligation is a HIPAA audit control, an FTC Safeguards Rule logging requirement, or an SEC recordkeeping rule. The specific words differ. The mechanism, general obligation narrowed silently in translation and never checked against the running system afterward, is identical, which is exactly why it belongs on this cross-cutting hub rather than being explained three separate times.

What a correct approach requires instead

The correct approach removes the translation gap by making the obligation and the check the same artefact, described in full on compliance is a specification problem. Instead of writing a policy sentence describing what the system should do and hoping engineering builds toward it, you write an assertion the system either satisfies or fails, derived directly from the regulatory text, with the regulatory citation attached to the assertion in the code.

The order matters as much as the content. A check written by looking at what the code already does and asserting that is what it should do proves nothing, because the standard was copied from the thing being measured rather than from the regulation. A check written first, from the obligation, before the implementation exists, is the only version that can actually fail, and a check that can fail is the only kind of check that tells you anything. That ordering is the entire difference between a check suite that is evidence and a check suite that is theatre.

The other structural piece the correct approach requires is that the thing grading the work is not the thing that did the work, covered in never let the model grade its own work. This applies whether the work was written by a person or generated: a reviewer with a stake in the outcome, whether that is pride of authorship or a shipped feature, is a worse judge of whether the requirement was actually met than an independent reviewer working from the same regulatory text.

How this plays out across a build's lifecycle

At design time, the naive approach produces a policy document. The correct approach produces an obligation list with the regulatory sentence quoted, not summarised, next to each item, split into what a person satisfies through process and what only the running system can satisfy, exactly as compliance is a specification problem describes. Only the second kind of item becomes engineering work, but it becomes engineering work immediately rather than after a compliance review flags it.

At implementation time, the naive approach lets engineers build the feature as scoped and trust that the policy, written separately, covers the compliance angle. The correct approach attaches the relevant assertions to the feature before code is written, so building the audit trail, the retention behaviour, or the human-review path is part of the ticket rather than a follow-up nobody schedules. Retention specifically deserves attention here, because it fails in both directions, keeping data too long and deleting it too early, across every copy the data has been replicated into, which retention and deletion across every copy covers in full.

At review time, the naive approach relies on a human reviewer remembering which regulatory provisions apply to the code in front of them, which is a lot to ask under a deadline. The correct approach lets the traced checks do that remembering: a change that removes an audit entry, widens data retention past its bound, or bypasses the data access layer described on building an audit-ready codebase fails a specific, named check rather than depending on a reviewer's memory of a rule they read months ago.

In production, the naive approach has no signal until an examination, an incident, or a complaint surfaces the gap between the policy and the system. The correct approach has the check suite running on every change, so a regression is caught the day it ships. Where a human is meant to review a system's output, when a human must stay in the loop covers the further failure mode where the review step exists on the org chart and not in substance, a reviewer who approves nearly everything because the design gives them no real basis to disagree.

A worked example, start to finish

Here is one obligation traced from regulatory text to production, because the pattern is easier to recognise in a single concrete case than as a general argument.

A financial services product stores customer records and is subject to a rule requiring secure disposal of customer information no later than a set period after last use. The naive approach writes a policy stating that customer data is deleted after that period, and implements a scheduled job that deletes rows from the customers table past the cutoff. The policy is accurate about intent and the job runs correctly against the table it targets. An auditor reading the policy and watching the job run would reasonably conclude the obligation is met.

What the policy does not mention, because nobody thought to enumerate it, is that the same customer data also lives in a support ticketing tool that synced a copy for every customer who ever contacted support, an analytics warehouse that ingested the full customer table for a dashboard project two years ago and was never told to expire anything, and a set of database backups taken nightly and retained for a year by a setting nobody has revisited since the infrastructure team configured it. The scheduled job deletes the row in the primary table. The other three copies keep the data for as long as their own, unrelated retention settings allow, which in the backup case is longer than the obligation permits and in the warehouse case has no bound at all.

Under the correct approach, the obligation is stated as "the customer's data is gone everywhere it was ever copied to," and that statement gets enumerated as a location inventory before any deletion mechanism is built, exactly as retention and deletion across every copy describes. The check that gets written asserts that a test record, run through the real deletion path, cannot be found in any of the four locations afterward. That bar matches what the regulation is asking for, rather than the narrower bar of confirming the scheduled job executed without error.

How to tell if your own team is at risk

A short set of direct questions surfaces whether an organisation is running the naive version.

Ask to see the obligation list with the regulatory text quoted next to each item, rather than a policy document summarising the regulation in the organisation's own words. If the list does not exist as a distinct artefact from the policy, the translation step that creates the gap has not been made visible to anyone.

Ask whether any compliance check in the codebase references the specific provision it exists to satisfy. If checks exist but nothing connects them to the regulatory text, coverage can only be assessed by inference, which is unreliable enough that almost nobody actually does it, and most teams assume they are fine rather than confirming it.

Ask, for a specific compliance check, whether it was written before or after the code it verifies, and check the commit history rather than accepting the answer from memory. A check written afterward encodes what the code already does and cannot fail, which means it was never testing anything.

And ask what a development partner would say if asked the same two questions we recommend asking any supplier: what share of the code is AI-generated, and what has to prove it works before it reaches your branch. If the second answer is vague, the partner's own process has the same gap this whole page describes, which is covered directly on choosing a development partner for regulated work.

Which obligations apply to your business is a question for counsel rather than for engineers, and nothing here is legal advice. What this guide does is take the obligations once somebody has scoped them and make them into properties a system can be tested against, continuously, by something that does not get tired at six on a Friday.

The document says what you decided. Only the running system decides what happens.

Explore the guide

Common questions

Why is AI-native development a better fit for regulated industries?

Because the main cost of compliance in software has always been writing and maintaining exhaustive checks, which was always given lower priority than visible features. When implementation is close to free that trade-off disappears, and the resulting check suite also serves as the evidence a regulator asks for, so the advantage is thoroughness rather than speed.

Is it safe to use AI-written code in a regulated system?

Only with independent verification. Veracode's spring 2026 testing found only 55 percent of model generations produced secure code while syntax correctness ran above 95 percent, so generated code compiles far more reliably than it protects, and in regulated systems that gap affects authentication, authorisation, and retention paths.

What makes compliance obligations easier to test than normal requirements?

Somebody already wrote them down in operative language and published them, which almost never happens in ordinary product work. A rule stating that an audit trail must record the identity of the individual deleting a record is nearly a test case, whereas a normal product requirement has to be discovered by releasing and arguing.

Why is a check suite better evidence than a policy document?

Because a policy describes intent at the moment it was written and does not change when the system stops matching it, while a check describes behaviour on every change. The run history is also credible precisely because it cannot be produced retroactively, which a document can be the week before an examination.

Does this replace legal advice or a compliance function?

No, and the boundary matters. A check suite cannot decide which obligations apply, interpret an ambiguous provision, or cover an obligation nobody translated into an assertion, so its coverage is exactly as good as the obligation list it was built from. Counsel scopes, the team translates, and the suite enforces from then on.

What should you ask a development partner about regulated work?

Ask what grades the code and whether it is the same context that wrote it, and ask whether the checks were written before or after the implementation. The order is visible in commit history, and a supplier who cannot show checks preceding the code they verify is describing a testing practice they do not have.

Where do compliance gaps usually appear in a working system?

In the parts nobody designed rather than the parts designed badly, which is why self-review finds so little. Logging built for writes misses reads, retention implemented in a delete button is bypassed by an export path, and each is the system working exactly as built.

Does the EU AI Act apply to software built outside the EU?

It can, because its scope follows where systems are placed on the market or used rather than where the developer is based, so a product sold into the EU may be in scope regardless of where it was built. Whether a specific product is covered, and in which risk category, is a determination for counsel.

Start a project