Software for compliance-heavy industries / Choosing a partner
Choosing a development partner for regulated work
Every development partner will tell you they take compliance seriously and that they test their code. Both statements cost nothing to make. This page covers the questions whose answers cannot be improvised, and what a good answer actually sounds like.
Published August 22, 2026. Editorial.
Key takeaways
- Ask what grades the code, and whether it is the same context that produced it. This separates practice from marketing faster than any other question.
- Ask whether checks were written before or after the implementation, then verify it in commit history.
- Ask what happens when the pipeline is failing and the release is due, because the answer describes their real process.
- Be suspicious of any supplier who claims compliance as an outcome rather than describing a mechanism.
Choosing a partner for regulated work is harder than choosing one for ordinary product work, because the thing that matters most is invisible in a portfolio. A supplier can show you a product that looks excellent and tells you nothing about whether its retention rules are enforced or its audit trail is complete.
So the assessment has to be about process, and process claims are cheap. These are the questions where a real practice and a described one give different answers.
"What grades the code?"
This is the most useful question, and it is worth asking exactly this way rather than asking whether they test.
The answer you want describes independence: the verification comes from a fresh context that did not produce the change, evaluating against criteria written separately, ideally with a different model doing the grading than did the work. The reasoning is on never let the model grade its own work, and it matters in regulated work more than anywhere.
The answer that should concern you is any version of "the agent writes the code and the tests and the pipeline passes." That is internal consistency, not verification. The tests encode the implementation's interpretation of the requirement, including whatever it misunderstood, so they pass while proving nothing about the requirement.
A supplier who has thought about this will recognise the question immediately and have a specific answer. A supplier who has not will answer a different question, usually about their test coverage percentage.
Listen also for what happens when you push past the first answer. A supplier might correctly say the grading is independent and stop there, satisfied that the word "independent" answered the question. A stronger answer explains what independence actually means in their setup: a separate run with its own context window and no access to the reasoning that produced the change, criteria pulled from a document rather than from the diff being graded, and a result that is recorded even when it fails rather than silently retried until it passes. That last detail matters more than it sounds like it should, because a grading step that quietly retries until it gets a passing result is not independent verification, it is the same process checking its own homework with extra steps.
"Were the checks written before or after the code?"
Order determines meaning. A check written after the implementation encodes what the implementation does, which is circular and in a compliance context worse than nothing, because it produces a passing result a reviewer will read as evidence.
The good answer is that checks derive from the specification before implementation exists. The useful follow-up is to ask to see it, because this claim is verifiable: in a repository, the commits show which came first. A supplier confident in the practice will be comfortable showing you. One who is not will explain why that is difficult to demonstrate.
There is a specific pattern in commit history that confirms the practice rather than merely being consistent with it. Look for a commit that adds a failing check with no corresponding implementation change, followed later by a commit that makes it pass. That sequence can only happen if the check was written first, because a check written after the fact is authored alongside the code that already satisfies it and never exists on its own in a failing state. A supplier who understands this distinction will be able to point to an example without having to search for one, because they will recognise it as the specific evidence the question is actually asking for.
"What happens when the pipeline is failing and the release is due?"
This question reveals the real process rather than the described one, and the answer is revealing regardless of what it is.
The concerning answers involve making the pipeline pass: skipping the check, adjusting the threshold, marking the test as flaky and moving on. Nobody says this in a sales conversation, so ask it as a scenario and listen for whether the answer includes any mechanism that prevents it.
Good answers include structural safeguards: changes to compliance checks are reviewed separately from the change that motivated them, the count of skipped checks is tracked and cannot rise without anyone noticing, threshold changes need separate approval, and removing a check traced to a provision requires a recorded written reason. These are the mechanics on using evals as compliance evidence, and a supplier who has them will describe them without prompting.
"How do you know an obligation is covered?"
The answer should involve a traceable mapping: an obligation list, assertions derived from it, and each check naming the provision it satisfies. Then coverage is a query.
If the answer is a coverage percentage, that is a different measurement answering a different question. Line coverage tells you what fraction of code was executed by tests. It tells you nothing about whether a regulatory obligation is enforced, and the two are frequently confused in supplier conversations because one of them is easy to produce.
"What do you do about generated code being insecure?"
A supplier working with AI-written code in a regulated context should be able to speak directly to this, because the evidence is public. Veracode's spring 2026 testing across more than 150 models found only 55 percent of generations produced secure code while syntax correctness ran above 95 percent [1].
A partner who acknowledges this and describes what they do about it is being honest with you. A partner who claims their process eliminates the issue, or who has not encountered the finding, is either not paying attention or is managing the conversation. Neither is what you want on a system holding customer records.
What to be suspicious of
Compliance claimed as an outcome. "Our platform is HIPAA compliant" is close to meaningless as stated, because compliance is a property of an organisation's whole operation, not a feature of a product. What a product can be is built to support specific obligations, described specifically.
Certification substituted for verification. A certification tells you an organisation passed an assessment at a point in time against a defined scope. It is worth having and it is not the same as the system enforcing your obligations, and a supplier who answers every technical question by referring to a certificate is avoiding the question.
Reluctance to show the repository. Most of the claims above are visible in a codebase. A partner unwilling to walk you through a representative one, under an appropriate agreement, is asking for more trust than the situation warrants.
A proposal that treats every regulated engagement the same way. Financial recordkeeping, healthcare privacy, and consumer credit obligations share a structure, which is the argument of this whole hub, but they are not identical, and a proposal that describes the same generic testing approach regardless of which one applies has probably not read your specific obligations yet. A partner who has actually scoped your situation will ask which regulator applies before pricing the work, not after.
Questions worth asking beyond the technical ones
The technical questions above tell you whether a supplier can verify their own work. A shorter set of questions tells you what happens when something is found wrong after the fact, which is a different and equally important property.
Ask who is accountable when an obligation turns out to be missed after release, and listen for whether the answer names a mechanism or a feeling. A supplier who says they will "work with you to fix it" has described an intention. A supplier who describes a specific process, how the gap is found, how it is triaged against severity, how affected records are identified and remediated, has described a mechanism, and the second is the one that will actually function under pressure.
Ask what happens to the check suite itself over the life of the engagement, since a suite that is not maintained decays the same way any other part of a codebase does. Obligations get amended, features get added that the original obligation list never anticipated, and a suite frozen at the moment of delivery slowly stops matching the system it is supposed to verify. A partner with an ongoing relationship to the codebase, rather than a one-time delivery, is in a better position to keep the two aligned.
What we would expect you to ask us
For symmetry, since this page would mean little otherwise: ask us the questions above and expect specifics. Ask what grades our code, and the answer is a fresh context with the criteria written from the specification first. Ask about the order, and the commits show it. Ask what happens when the pipeline is failing before a release, and the answer is that the check suite changes go through separate review, because the alternative is a control that the people it constrains can disable without anyone noticing.
We would also expect you to ask what we do not do. We do not determine which obligations apply to your business, and we would want counsel involved in that before we write a single assertion, because a rigorously verified implementation of the wrong requirements is an expensive way to be non-compliant.
The broader question of selecting a partner is covered in choosing a software development partner. If you want to talk about a regulated build, get in touch.
Best for
- Buyers evaluating suppliers for a build in a regulated sector
- Teams who have been offered a proposal and want to test the process claims in it
Avoid if
- The obligations have not been scoped yet, in which case that work comes before supplier selection
Check before you decide
- Ask what grades the code and whether it is the same context that wrote it
- Ask to see commit history showing checks written before implementations
- Ask what happens when the pipeline is failing and the release is due, and listen for a mechanism
- Ask how they know a specific obligation is covered, and reject a coverage percentage as the answer
Common questions
What is the single best question to ask a development partner about regulated work?
What grades the code, and is it the same context that produced it. A supplier with a real practice recognises the question and answers it specifically, while one without it answers a different question, usually about test coverage percentage.
How can you verify that checks were written before the implementation?
In the commit history, since the order is visible in any repository. A supplier confident in the practice will be comfortable showing you, and one who is not will explain why demonstrating it is difficult.
Why ask what happens when the pipeline is failing before a release?
Because it shows the real process rather than the described one. The concerning answers involve making the pipeline pass by skipping a check or adjusting a threshold, and good answers include structural safeguards like separate review of check changes and a tracked count of skipped checks that cannot rise without anyone noticing.
Is test coverage a useful measure of compliance?
No, and the two are frequently confused because one is easy to produce. Line coverage says what fraction of code was executed by tests and says nothing about whether a regulatory obligation is enforced, which is answered by tracing each check to the provision it satisfies.
What claims should make a buyer suspicious?
Compliance claimed as a product outcome rather than described as support for specific obligations, certification offered as an answer to technical questions about enforcement, and reluctance to walk through a representative repository under an appropriate agreement. Most process claims are visible in a codebase, so unwillingness to show one asks for more trust than the situation warrants.
What does it mean for a supplier to say their product is compliant?
Taken literally, little, because compliance describes a property of an organisation's whole operation rather than a feature of a product. A more useful and honest claim is that a product is built to support specific named obligations, described one at a time, which a buyer can then check against the obligation list rather than accepting as a general label.
How should a buyer run a supplier evaluation for a regulated build?
Ask the five verifiable questions directly rather than accepting a general assurance: what grades the code and whether it is the same context that wrote it, whether checks were written before or after the implementation with commit history to show it, what happens when the pipeline is failing before a release, how obligation coverage is known, and what the supplier does about generated code being insecure by default.
What is the risk of hiring a partner who cannot show checks preceding code?
The risk is that verification is circular: a check written after an implementation encodes what that implementation already does, so it passes while proving nothing about the underlying obligation. In a regulated system this produces a passing pipeline that a reviewer or examiner will reasonably treat as evidence, when it demonstrates only that the code is internally consistent with itself.
How a build like this runs
Related reading
Red flags when hiring a software development agency
The wrong agency costs you months, not just money. Here are the warning signs worth checking before you sign a contract, not after.
How to choose the right external development team
Hiring an outside team is a decision with serious consequences. The wrong partner costs you time and progress you cannot get back.
Never let the model grade its own work
When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.
Software development partner vs. vendor: what is the real difference?
A vendor builds what you ask for. A partner tells you when what you asked for is wrong. The difference shows up months after the contract is signed.