Building compliant software in healthcare / Auditing what you have
Where PHI leaks in a working application
When a healthcare application has a PHI problem, it is rarely in the clinical database. That is the part everyone thought about. The problem is in the eleven other places clinical data was stored, each created for a good reason by someone solving a different problem.
Published August 22, 2026. Editorial.
Key takeaways
- Enumerate every location where PHI is stored before assessing any control, because a control you cannot scope is a control you cannot evaluate.
- Logs, error reports, analytics, and support tooling are the most common unintended stores.
- Non-production environments filled with production data are a recurring and under-examined exposure.
- Deletion almost never reaches the other copies, so records removed from the primary store persist in the copies indefinitely.
Ask a healthcare engineering team where PHI is stored and the first answer is the clinical database. It is usually correct about that database: encrypted, access-controlled, logged, backed up, thought about carefully by people who understood the obligation.
The list is longer. Here is what a full enumeration typically finds, roughly in order of how often it surprises the team.
Application logs
This is the most common place. PHI reaches logs in three ways, none of them deliberate.
Error handlers that capture request context, including bodies. A clinical endpoint receives patient data, an exception occurs, and the payload is written to the log with the stack trace.
Debug logging left in place. Someone adds a log line while investigating an issue, it prints a record for context, and it is released. Log lines are rarely removed once added, and nobody reviews them for content.
Query logging at the database or ORM layer, which writes the parameters, and the parameters are patient identifiers and record content.
Then the log is sent to an aggregation service, often third-party, and PHI has left your infrastructure. The check is direct: take a production log sample and read it, searching for identifiers and record content. Most teams have not done this, and doing it once is usually informative.
Error tracking and observability
Related but distinct, because these services are designed to capture rich context. Exception reports include local variables, request payloads, user context, and a record of the user's recent actions. Session replay tools record the rendered screen, which on a clinical product means the chart.
Distributed tracing captures span attributes, and if a span records a query or a record identifier, that is PHI in a trace store with its own retention and access model, usually configured by whoever set up tracing.
Analytics
Product analytics events carry properties, and properties carry whatever the developer attached. A user identifier that can be linked to a patient is PHI. An event property containing a condition, a medication, or a department is PHI. Form field values captured by an auto-instrumenting analytics library are PHI.
Analytics platforms typically have long retention by default, broad internal access, and no business associate agreement, which makes this one of the higher-risk categories.
Support tooling
Support platforms sync customer or patient context so agents have background. Ticket bodies contain whatever the patient or the staff member typed, which in healthcare is frequently clinical detail. Attachments contain documents. This store is accessed by a broad group of people, and its access controls are the vendor's rather than yours.
Non-production environments
Staging and development databases filled with copies of production data, because realistic data makes testing better and anonymised data is work. This is common, and it is one of the largest exposures in a typical healthcare stack, because non-production environments have weaker access controls, more people with access, less monitoring, and frequently no encryption on their backups.
The related pattern is a production data sample copied to a developer machine to reproduce a bug, which then stays in a downloads folder indefinitely.
Exports and reports
Scheduled exports to object storage for reporting or partner delivery. Ad hoc CSVs generated for a meeting. Emailed reports. Files generated by a batch job and left in a bucket whose lifecycle policy was set once and never revisited.
Exports deserve particular attention because they place a copy outside every control the application has, and because the bucket policy is usually configured by whoever built the export rather than by anyone thinking about retention.
Caches, queues, and search
A cache holding record content. A message queue with payloads in transit, which are at rest while queued. A search index holding denormalised copies of record text, often with its own access model and no audit logging at all.
Search indexes are the most consistently overlooked, because teams think of them as derived data rather than as a copy.
Backups, including the old ones
Backups of everything above, not just the primary database. Snapshots taken before a policy changed. Backups in a different cloud account or region with different controls. Long-retention archival tiers holding data that should have been disposed of years ago.
AI feature data
If the product has AI features, add prompt logs, completion logs, prompt caches, evaluation datasets assembled from production traffic, and traces. These are covered on building AI features on clinical data, and they are new enough that they are rarely in anyone's inventory.
How to actually find them
The useful exercise is to check this list against your own system, and there are three methods that find different things.
Follow the data outward from the clinical record. Take one patient record and trace every place a copy of any part of it could be stored. Reads, writes, replication, caching, indexing, logging, analytics, export, backup.
Follow the network. Enumerate every external destination the application transmits to, from configuration and from observed traffic. Each one either receives PHI or does not, and the honest answer for several will be "we would have to check."
Follow the people. Ask each team what data they work with and where they get it. Support, analytics, data science, and QA all have working copies, and they are frequently in places engineering did not provision.
Then check deletion
Once the list exists, apply the test that finds the most: take a record that should have been deleted and search for it everywhere on the list.
Deletion almost never reaches the other copies. It is built as a database operation, sometimes with a cascade to related tables, and it stops there. The record persists in the search index, the analytics platform, the support tool, the export bucket, the backups, the staging environment, and the trace store. Every one of those is a copy of data somebody asked to have deleted.
Building deletion that reaches every copy is real work and it is worth scoping deliberately rather than discovering during an incident. The general audit process is on how to audit an existing financial application, and it applies directly.
If you want this run against your product, get in touch.
Best for
- Healthcare products in production for more than a year, especially with a mature observability stack
- Teams preparing for a security risk analysis, a certification, or enterprise procurement
Avoid if
- The product has not launched and holds no real clinical data, where building the controls in is cheaper than mapping later
Check before you decide
- Read a real production log sample and search it for identifiers and record content
- List non-production environments and check whether any is filled with production data
- Take a deleted record and search for it across every store on the list
- Check whether the search index has any audit logging at all
Common questions
Where does PHI most often end up outside the clinical database?
In application logs, error tracking and observability tooling, product analytics, support platforms, non-production environments filled with production data, exports, caches and search indexes, and backups of all of the above. Each was created for a good reason by someone solving a different problem, which is why none of them was assessed as a clinical data store.
How does PHI get into application logs?
Through error handlers that capture request bodies alongside stack traces, debug log lines added during an investigation and never removed, and query logging that writes parameters containing identifiers and record content. The direct check is to take a production log sample and read it, which most teams have never done.
Are staging and development environments a real exposure?
They are one of the largest in a typical healthcare stack, because filling them with production data is common and anonymising the data is work nobody scheduled. Non-production environments have weaker access controls, more people with access, less monitoring, and often no encryption on their backups.
Why are search indexes so often missed?
Because teams think of them as derived data rather than as a copy of the record. An index typically holds denormalised record text, has its own access model, and frequently has no audit logging at all, which makes it a PHI store with fewer controls than anything else in the stack.
What single test finds the most problems?
Take a record that should have been deleted and search for it across every store on your list. Deletion is almost always built as a database operation with a cascade to related tables and stops there, so the record persists in the search index, analytics, support tooling, exports, backups, staging, and traces.
Do AI features add new places PHI can leak?
Yes, and they are new enough that they are rarely in anyone's inventory. Prompt logs, completion logs, prompt caches, and evaluation datasets assembled from production traffic all hold PHI once a clinical prompt or output contains patient information, and they are typically outside the access controls, retention rules, and audit logging applied to the rest of the system.
How should a team methodically find where PHI is stored?
Three methods find different things: follow the data outward from one patient record through every read, write, cache, and export it touches; follow the network by enumerating every external destination the application transmits to; and follow the people by asking each team, including support and data science, what data they actually work with and where they got it.
Are exports and backups worth auditing separately from live systems?
Yes, because they place a copy of PHI outside every control the application enforces. Ad hoc exports are often configured once by whoever built them and never revisited, and backups make the problem worse by preserving old snapshots taken before a retention or encryption policy changed, sometimes in a different cloud account with weaker controls.
Related reading
A practical pre-launch security review for a small team
You do not need perfect security to launch. You need to check the few basics that find most real problems, and to know when the risk is big enough to bring in a specialist.
The early warning signs your software project is in trouble
A software project rarely fails in one dramatic moment. It falls behind slowly, and the early signs are easy to explain away. Here are the ones to watch and what to do about each before it is too late.
What technical debt really is, and when to pay it back
Technical debt is not messy code. It is a deliberate choice to give up some quality for speed. Here is how to tell smart debt from reckless debt, and when to pay each one back.