Building it right

Building AI features on clinical data

Adding a model to a clinical product raises three questions most teams answer implicitly: where the data goes, whether the prompts and outputs are records, and whether the feature has crossed from software into something regulated as a medical device. Answering them explicitly, before building, is much cheaper than answering them afterwards.

Published August 22, 2026. Editorial.

Key takeaways

  • A model provider receiving PHI is a business associate like any other vendor, and default consumer terms are not sufficient.
  • Prompts and outputs containing PHI are PHI, which means they need the same logging, retention, and access treatment as anything else.
  • There is a regulatory boundary between software that informs a clinician and software that makes a clinical determination, and it is worth finding early.
  • Evaluation matters more here than in any other AI application, because a plausible wrong answer about a patient is the failure mode.

Clinical AI features get built the same way other AI features get built: someone prototypes with an API key, it works impressively, and the question of what happens to the data arrives later. In healthcare, later is expensive.

Three questions are worth answering before the prototype becomes a product.

Where does the data go?

A model provider that receives PHI is a business associate. That is the same analysis as any other vendor, covered on business associates and third-party services, and the practical consequences are specific.

Default consumer or developer terms are typically not sufficient for PHI. Providers generally offer distinct enterprise arrangements with different data handling, retention, and training commitments, and the difference between those tiers is exactly the difference that matters here. An API key created in five minutes on a standard plan is almost certainly the wrong basis for a clinical feature.

The questions to resolve are concrete. Is the data used for training? What is the retention period on prompts and completions, and can it be set to zero? Where is inference performed geographically? Is an agreement in place covering PHI? What subprocessors are involved?

None of these is unanswerable, and all of them are much easier to resolve before a feature is released than after.

Prompts and outputs are records

This is the part teams miss most often, and it follows directly from the definition.

If a prompt contains patient information, that prompt is PHI. If a model output describes a patient, that output is PHI. Which means both are subject to the same treatment as any other PHI in your system: access controls, audit logging, retention rules, and inclusion in the inventory of where PHI is stored.

In practice, prompt and completion data usually ends up somewhere nobody classified. A logging pipeline capturing requests and responses for debugging. A prompt cache. An evaluation dataset assembled from production traffic, which is enormously useful for improving the feature and is a PHI store that frequently has none of the controls the clinical database has. A trace stored in an observability platform.

Every one of those is a legitimate engineering need, and every one creates a copy of clinical data outside the paths the compliance work covered. The answer is not to avoid them; it is to classify them as PHI stores from the start and apply the same controls, which is straightforward when done at design time and awkward when discovered in an audit.

Is your feature now a regulated medical device?

There is a regulatory boundary between software that helps a clinician and software that makes a clinical determination, and which side a given feature is on is a genuine legal question that depends on what the software does, what it claims, and how much the clinician can independently review its basis.

We are not going to try to define that boundary here, because it depends on specifics and it is a determination for regulatory counsel. What we will say is that the boundary exists, that it is much cheaper to find before building than after, and that the framing of a feature affects which side it falls on.

The practical version for a product team: if a feature is moving towards recommending a diagnosis, a treatment, a dosage, or a triage priority, get a regulatory opinion before the build rather than before the launch. Features that summarise, retrieve, draft, and organise carry much less regulatory risk, and a great deal of clinical value is found in those features.

Evaluation matters more here

Everything we say about evaluation in eval-driven development applies more strongly to clinical features, for a specific reason.

The characteristic failure of a language model is output that is plausible and wrong. In most products, plausible and wrong costs a user some time. In a clinical context, a fluent, confident, incorrect statement about a patient is the exact kind of harm the whole regulatory apparatus exists to prevent, and it is far more dangerous than an obvious error because it does not trigger scepticism.

Three things follow for a build.

The evaluation set has to include the hard cases, deliberately. Not the representative sample, which the model handles well. The rare presentations, the ambiguous records, the contradictory histories, the cases with missing data. Performance on the common case is not informative about safety.

The model must not grade its own output. A model asked to check its own clinical summary will confirm it, because the reasoning that produced the summary also produces the confirmation. The grading has to come from a fresh context against criteria written independently, and for clinical content, from a qualified human on a sampled basis. We covered the mechanism in never let the model grade its own work.

The feature has to show its uncertainty. A model asked to summarise a record with a contradiction in it will typically produce a clean summary that silently picks one side. The design requirement is that the feature says when it is unsure, and that the interface makes the underlying source reachable so a clinician can check rather than trust.

Design so the clinician can verify

The single most useful design property for a clinical AI feature is that a person can check its work quickly.

That means output that cites the specific record, note, or result it drew from, with a link to it. It means showing what the model saw, not only what it concluded. It means making disagreement easy: a one-click path to the source beats any amount of confidence calibration.

This is also the property that keeps a feature on the lower-risk side of the regulatory boundary discussed above, because software whose basis a clinician can independently review is in a materially different position from software asking to be trusted.

If you are scoping a clinical AI feature and want the data-flow and evaluation questions worked through before the build, get in touch. Our AI development work covers the general case, and building AI products covers the path from demo to production.

Best for

  • Summarisation, retrieval, drafting, and organisation features where a clinician reviews the output
  • Teams willing to build the evaluation set before the feature rather than after

Avoid if

  • The feature is moving towards a diagnosis, treatment, dosage, or triage determination and no regulatory opinion has been obtained
  • The only available model access is on default consumer terms

Check before you decide

  • Confirm the provider tier, retention setting, training commitment, and agreement before any PHI is sent
  • List every place prompts and completions are stored, including logs, caches, traces, and evaluation datasets
  • Check that the evaluation set contains the hard cases, not a representative sample
  • Check that the feature cites its sources and makes them reachable in one step

Common questions

Can we send PHI to a commercial model provider?

Only under terms appropriate for PHI, since a provider receiving it is a business associate like any other vendor. Default consumer or developer terms are typically insufficient, and the questions to settle first are training use, prompt and completion retention, inference location, the agreement itself, and subprocessors.

Are prompts and model outputs subject to HIPAA?

If a prompt contains patient information it is PHI, and if an output describes a patient it is PHI, which means both need the same access controls, audit logging, retention rules, and inventory treatment as anything else. This matters because prompts and completions are usually stored in logging pipelines, caches, traces, and evaluation datasets that nobody classified as clinical stores.

When does a clinical AI feature become a regulated medical device?

The determination depends on what the software does, what it claims, and whether a clinician can independently review the basis for its output, and it is a decision for regulatory counsel rather than an engineering judgment. The practical rule is to get an opinion before building anything moving towards a diagnosis, treatment, dosage, or triage recommendation, since summarisation, retrieval, and drafting carry far less regulatory risk.

Why does evaluation matter more for clinical AI features?

Because the characteristic model failure is output that is plausible and wrong, and a fluent confident incorrect statement about a patient is more dangerous than an obvious error, since it does not trigger scepticism. The evaluation set therefore has to be built from hard cases, rare presentations, ambiguous records, and missing data rather than a representative sample that the model already handles well.

What design property makes a clinical AI feature safest?

That a clinician can check its work in one step. Output that cites the specific note or result it drew from, with a link, and that shows what the model saw rather than only what it concluded, both reduces harm and keeps the feature on the lower-risk side of the regulatory boundary, because software whose basis can be independently reviewed is in a different position from software asking to be trusted.

Is it safe to build a clinical AI feature at all?

Safety depends on what the feature does. Summarisation, retrieval, drafting, and organisation features that a clinician reviews carry less regulatory risk once the data-flow and evaluation questions are answered before building. Features moving toward a diagnosis, treatment, dosage, or triage recommendation carry real regulatory risk and need a legal opinion before the build starts, not before launch.

Why resolve data-flow and evaluation questions before building, not after?

Because a prototype that works impressively tends to become the released feature by default, and retrofitting a business associate agreement, retention settings, and an evaluation set onto a feature already in clinical use is far more expensive than deciding them at the outset. The technical work of adding the model is usually the smallest part of the effort.

How is a summarisation AI feature different from a diagnostic one under compliance?

A summarisation, retrieval, or drafting feature that a clinician reviews is far from the boundary between software that informs and software that makes a clinical determination. A feature recommending a diagnosis, treatment, dosage, or triage priority is close to or across that boundary, and where a given feature falls depends on what it does and claims, which is a determination for regulatory counsel rather than engineering judgment.

Start a project