Software for compliance-heavy industries / Rules that cross every sector
Retention and deletion across every copy
Every regulated industry has retention obligations, and they fail the same way in every one of them. The policy is enforced on the primary database, which is the one place somebody thought about, and ignored in the eleven other places the data was stored.
Published August 22, 2026. Editorial.
Key takeaways
- Retention fails in both directions: data kept past its disposal date, and data deleted automatically before its preservation period ends.
- Enumerate every location where data is stored before assessing any retention control.
- Deletion is almost always built as a database operation and almost never reaches the other copies.
- Backups are the hardest case and need a deliberate, written position rather than an assumption.
Retention obligations appear in every regulated sector, phrased differently and describing the same engineering problem. The FTC Safeguards Rule requires secure disposal of customer information no later than two years after last use, with exceptions, and periodic review of the retention policy to minimise unnecessary retention [1]. Recordkeeping rules require preservation for defined periods [2]. Privacy frameworks give individuals rights to have data deleted.
All of them assume you know where the data is. That assumption is where the failures start.
Both directions are failures
Teams think of retention as a keeping-too-long problem. It is equally a deleting-too-early problem, and the second is harder to notice.
Keeping too long: a record passes its disposal date and remains in an export bucket, a backup tier, a support tool, and a warehouse. Nobody notices because nothing breaks.
Deleting too early: a record subject to a multi-year preservation obligation is stored in a system whose retention was configured for cost by whoever set it up, with a ninety-day default. The record is gone, the obligation is unmet, and nobody finds out until somebody asks for it. This is especially common for audit logs and system records, which are frequently stored in observability tooling whose retention was never aligned with the obligation that depends on it.
A concrete version of the second failure: a brokerage's transaction logs sit in a logging platform configured with a ninety-day default because that was the setting when someone signed up for the tool, years before anybody connected that platform to the recordkeeping rule requiring years of preservation. The gap is invisible for the entire ninety days, because the logs exist and look complete. It becomes visible only when an examiner asks for a record from four months earlier and the honest answer is that it no longer exists. Nothing about that failure resembles a security incident. Nothing alerted, nothing crashed, and the dashboard that would have shown a retention setting misaligned with an obligation was never built, because retention configuration is normally treated as an infrastructure decision rather than a compliance one.
Check both. A retention control that only prevents over-retention is half a control.
Enumerate before you assess
You cannot evaluate a retention control without knowing what it is supposed to cover, so the enumeration comes first. The full method is on the healthcare hub at where PHI leaks in a working application, and the list generalises to any sector.
The locations that recur: primary database, read replicas, analytics warehouse, backup snapshots across tiers, object storage holding exports and attachments, search indexes, caches, message queues, application logs, error tracking, product analytics, support tooling, email systems that received a record in a template, non-production environments filled with production data, local copies on developer machines, and, for products with AI features, prompt logs, completion logs, evaluation datasets, and traces.
For each, record what it holds, its current retention, who can reach it, and which obligation governs it. That table is the deliverable, and building it is usually the point at which somebody says "I did not know we had that."
The enumeration is genuinely harder than it sounds, and the difficulty is worth naming rather than glossing over. A platform team can list the systems it operates directly. It usually cannot list every third-party tool an employee connected to production data over several years, because that connection often required no approval: an API key pasted into a support tool, a webhook forwarding customer events into an analytics product, a spreadsheet export somebody set up once for a board meeting and never disabled. None of those show up in an infrastructure inventory, because none of them are infrastructure the platform team provisioned. The practical starting point is not a technical scan but a short, direct question put to every team that touches customer data: what have you connected, and what does it store. The answers are usually incomplete on the first pass and improve on the second, once people start noticing connections they had stopped thinking about.
Deletion does not reach the copies
Here is the test that finds the most in the least time. Take a record that should have been deleted and search for it in every location on your list.
Deletion is almost universally built as a database operation, sometimes with a cascade to related tables, and it stops at the edge of the database. The record persists in the search index, the analytics platform, the support tool, the export bucket, the backups, the staging environment, and the trace store. Every one of those is a copy of data somebody asked to have deleted, or that an obligation required to be disposed of.
Building deletion that reaches every copy is real work and it is worth scoping deliberately. The pattern that works is a deletion event that is sent to every store, with each store's handler acknowledging completion, and a reconciliation process that verifies the record is actually gone rather than trusting that the handler ran. Without the verification step, a handler that silently fails leaves data in place and reports success, which is worse than not having the mechanism, because now you believe it worked.
The edge case that trips up an otherwise sound design is the derived copy: a value computed from the deleted record rather than a duplicate of it. A customer's spending pattern, computed nightly into an aggregate table used for fraud scoring, contains no field that names the customer directly, so a search for the customer's identifier does not find it. But the aggregate was built from that customer's transactions, and deleting the transactions without recomputing or removing the aggregate leaves a trace of the person behind in a form the enumeration will miss unless someone thought specifically about derived data as its own category. The fix is to treat any table computed from regulated data as itself governed by the same obligation, and to include a recomputation or removal step in the deletion event's list of handlers, not just a search for direct copies.
Backups are the genuinely hard case
Backups exist to be immutable and complete, which directly conflicts with selectively removing records from them. Rewriting historical backups to remove a record is usually impractical and sometimes defeats the purpose of having them.
There is no simple technical answer, and any guide that offers one is oversimplifying. What organisations do in practice is take a written position: backups have a defined maximum retention, deletions are applied to the live system immediately, and any restore from backup runs a reconciliation that re-applies pending deletions before the restored data is used. That position is defensible when it is written down, bounded, and actually implemented, particularly the restore reconciliation, which is the part that usually exists only in theory.
The failure is not having a considered position. It is having never thought about it, so backups retain everything forever, without anyone noticing, with no bound and no reconciliation, and nobody can articulate why that is acceptable.
Non-production is where this causes the most problems
Staging and development environments filled with production data are common, and they are the place the retention control does not reach. A deletion applied to production does not reach the copy in staging. A record disposed of on schedule in the live system persists in a development database somebody restored eight months ago.
The simplest answer is not to have production data in non-production environments, using generated or properly de-identified data instead. That is more work up front and it removes an entire category of problem, including the access control and encryption problems that come with the same copies. Where it is genuinely impractical, the fallback is a bounded refresh cycle where non-production data is destroyed and refilled on a schedule, so the maximum age of a copy is known.
Properly de-identified is doing real work in that sentence and deserves a separate warning, because a copy that looks de-identified is not the same as one that is. Removing a name field and leaving a date of birth, a zip code, and a rare combination of purchase history intact can still identify a specific person once those fields are combined, even though no single field on its own looks like a name. A staging environment built by copying production and running a script that blanks the obvious identifying columns is a common shortcut, and it is also a common way to end up with a copy that is regulated data wearing a disguise, subject to the same retention obligation as the record it was copied from, while nobody treats it that way because it no longer has a name attached.
Make it a check
Everything above can be asserted rather than reviewed annually.
Assert that each store's configured retention matches the obligation that governs it, which catches the day somebody reduces a log retention for cost. Assert that a deleted record is absent from every store, using a test record run through the real deletion path. Assert that no non-production environment holds data older than the refresh bound. Assert that the reconciliation process ran and reported zero outstanding deletions.
Those four turn retention from a policy that is true when written into a property that is true continuously, which is the whole argument of this hub applied to one obligation.
If you want help building the inventory, get in touch.
Best for
- Products with a mature data stack where copies have accumulated over several years
- Teams facing privacy requests they currently satisfy only in the primary database
Avoid if
- The applicable retention periods have not been determined, since both over- and under-retention are judged against them
Check before you decide
- Take a deleted record and search for it in every store you can name
- Compare each store's configured retention against the obligation governing it, in both directions
- Ask what happens to pending deletions when a backup is restored
- Ask whether any non-production environment holds production data, and how old the oldest copy is
Common questions
What are the two ways retention controls fail?
Keeping data past its disposal date, and deleting data automatically before a preservation period ends. The second is harder to notice and common for audit logs stored in observability tooling whose retention was set for cost by whoever configured it, with no reference to the obligation depending on it.
Why does deletion rarely reach every copy?
Because it is built as a database operation with a cascade to related tables and stops at the edge of the database. The record then persists in the search index, analytics platform, support tooling, export buckets, backups, staging environments, and trace stores, none of which the delete path was ever built to reach.
How should deletion across every copy be built?
As a deletion event that is sent to every store, with each handler acknowledging completion, plus a reconciliation process that verifies the record is actually gone rather than trusting the handler ran. Without verification a silently failing handler leaves data in place and reports success, which is worse than having no mechanism because you now believe it worked.
What do you do about deletions and backups?
Take a written position rather than pretending there is a simple technical answer: bounded maximum backup retention, immediate deletion in the live system, and a restore reconciliation that re-applies pending deletions before restored data is used. That is defensible when written down and actually implemented, and the restore reconciliation is the part that usually exists only in theory.
Should non-production environments hold production data?
Preferably not, since generated or properly de-identified data removes an entire category of retention, access control, and encryption problems at once. Where that is impractical, bound the exposure with a scheduled refresh cycle so the maximum age of any copy is known rather than open-ended.
What does a retention inventory actually contain?
A table listing every location where data is stored, including the primary database, replicas, backups, warehouses, search indexes, logs, and any prompt or completion logs from AI features, with each row recording what the location holds, its current retention setting, who can reach it, and which obligation governs it. Building the table is usually the point at which a team discovers a copy it did not know existed.
How should a deletion request be tested?
By taking a real test record, running it through the actual deletion path, and then searching for it in every location on the retention inventory rather than trusting that the delete operation succeeded. This finds the gap directly: deletion is almost always built as a database operation, so the record commonly still exists in the search index, exports, backups, or a support tool.
How long does it take to fix retention across every copy?
Fixing configured retention settings is usually the fastest step once the inventory exists, often a matter of configuration changes and a storage cost conversation. Building deletion that is sent to every store with verified reconciliation is a larger undertaking and worth scoping deliberately rather than attempting all at once, since the enumeration itself determines how much work remains.
References
- [1] 16 CFR 314.4(c)(6): secure disposal of customer information no later than two years after last use, subject to stated exceptions, and periodic review of the retention policy.
- [2] 17 CFR 240.17a-4(f)(2)(i)(A): preservation of a record for the duration of its applicable retention period with a complete audit trail.
Related reading
A practical pre-launch security review for a small team
You do not need perfect security to launch. You need to check the few basics that find most real problems, and to know when the risk is big enough to bring in a specialist.
What technical debt really is, and when to pay it back
Technical debt is not messy code. It is a deliberate choice to give up some quality for speed. Here is how to tell smart debt from reckless debt, and when to pay each one back.