Auditing what you have

Compliance gaps in proptech products

Real estate products accumulate a characteristic set of gaps, and they differ from the ones in finance and healthcare. The problem is rarely a missing control around the data. It is that nobody can say what the software decides, who it decided it about, or why.

Published August 22, 2026. Editorial.

Key takeaways

  • The most common gap is that a feature became a decision maker and no review was triggered.
  • Most products cannot report their own outcome distribution, which means they cannot answer the first question anyone will ask.
  • Ranking, feeds, and notifications are almost always excluded from the compliance conversation despite deciding who sees what.
  • Explanations that are reconstructed rather than captured fail at the first real challenge.

Auditing a proptech product is a different exercise from auditing a financial or healthcare application. In those industries you are mostly checking controls around data: is it logged, is it encrypted, is it retained correctly. Here you are checking what the software decides and whether anyone can account for it.

These are the gaps that recur.

The undeclared transition

The most common finding is that a feature crossed from informing a decision to making one, and nothing in the organisation noticed the change.

It shows up in specific ways. A score that customers were meant to interpret is now connected to an automatic action. A recommendation that was advisory is followed in the overwhelming majority of cases, which makes it effectively determinative even though a human is technically still involved. A threshold configured by a customer, in a settings page nobody on the product team has looked at in a year, that auto-declines below a value.

The audit question is simple and rarely asked: for each output the system produces, what happens next, and how often does a human actually change it? A human review step that overturns the model in a negligible fraction of cases is not functioning as a review, whatever the process diagram says. That is worth knowing, because the compliance position depends on what happens rather than on what is described.

A common version of this pattern: a screening product markets itself, correctly at launch, as decision support with mandatory human review. Two years later, an enterprise customer has configured the workflow so that a specific staff member is nominally the reviewer of record but, by the customer's own account, approves whatever the score recommends because the volume of applications makes individual review impractical. Nothing in the software changed. The vendor's marketing copy, written at launch, is still technically accurate about what the product supports. But for this specific customer's specific deployment, the product is making the decision, and the vendor's own compliance documentation, if it exists at all, was written about the version of the product that no longer describes how this customer actually uses it. Catching this requires looking at how customers actually configure the product, not at what the product's default configuration or marketing materials say it does.

The product cannot report its own outcomes

Ask a proptech team for the distribution of outcomes across groups for the last quarter and the usual answer is that it would take some work.

That answer is itself the finding. If the product cannot report this, then nobody, including the team, knows whether the system produces disparate outcomes, and no assurance offered about fairness is based on evidence. It also means the question cannot be answered when it arrives from outside, at which point the work happens under much worse conditions.

The underlying cause is usually a data model decision made early: outcomes were recorded for operational purposes rather than for analysis, group information was never collected or was collected in a form that cannot be joined, and decisions were logged in a way that captures the result but not the context. Fixing it retrospectively is real work, and it only produces data from the point of the fix forward, which is why finding this early matters more than most findings.

Ranking was never in scope

Almost every proptech compliance conversation covers screening or valuation and stops there. The ranking model, the recommendation feed, the notification logic, and the search ordering are treated as product features rather than as systems deciding who learns about which housing opportunities.

The test is to ask who has ever reviewed the ranking system from a fair housing perspective. In most organisations the answer is nobody, and the ranking team has never been in a meeting where the question came up. That is not negligence; it is a scoping failure that follows from the way the compliance conversation is usually framed. The reasoning is on housing advertising and audience targeting.

Explanations are reconstructed

Ask how a decision from six months ago would be explained. Watch what happens.

If the answer involves re-running the model, or looking at the current record, or a data scientist investigating, the product cannot explain historical decisions. It can produce a plausible account, which is different and worse, because it will sometimes disagree with what actually happened.

The specific things to check are whether factor attribution was computed at decision time, whether inputs were captured as values rather than as references, and whether the model version is identified and the artefact retained. Most products fail at least one, and the model retention one is the most commonly missed because retaining old model artefacts feels like clutter until the moment it is the only thing that would have helped. Covered on explainability and adverse action.

Third-party data has no provenance

Screening and valuation both depend on purchased data, and products routinely store the values without recording where each came from or when.

The consequence appears when someone disputes an item. Without provenance you cannot tell which source supplied it, cannot tell whether the item has since been corrected, and cannot identify the other decisions that used the same bad data. The dispute gets handled for the one person who complained, and the same erroneous data continues to affect everyone else it touched.

A specific version of this to check for directly: an eviction filing that was later dismissed, withdrawn, or resolved in the tenant's favour, but that continues to appear on new screening reports because the screening company's database, or the product's own cache of it, was never updated to reflect the outcome. If the product does not record when each data element was retrieved and from which specific query or feed, it cannot distinguish a report pulled before the correction from one pulled after, which means it cannot tell whether the current data is even the corrected version. Testing this concretely means picking a handful of applicants whose records are known to have changed since they were screened, and checking whether the product's stored decision reflects the old data or the new. If the answer is that nobody knows without manually checking, the provenance gap is real rather than theoretical.

The vendor model nobody can explain

Products that use a third-party score often cannot describe how it works, and the vendor treats the model as proprietary.

That is a commercially normal arrangement with an awkward consequence: when a decision is challenged, the answer is that a vendor's model produced a number. Whether that is adequate depends on the framework and the context, and it is a question for counsel. What is clear is that it is a weaker position than being able to explain the decision, and that it should be a considered choice rather than something discovered when the first challenge arrives.

The audit question here is narrower than "is the vendor's model good," which is difficult to answer from outside and is not the point. It is whether the contract with the vendor gives the provider anything to work with when a decision is questioned: a right to request the specific factors behind an individual score, a commitment from the vendor to cooperate with a dispute, or at minimum a description of the categories of data the score draws on. A contract that gives the provider no way to see how the score was reached, with no such provision, leaves the provider with nothing to offer an applicant beyond the number itself, and that gap is worth finding during procurement rather than during a live dispute, when the ability to negotiate a better term is gone.

How to run the review

The order that works follows the general approach on how to audit an existing financial application, with a different emphasis.

Start by listing every output the system produces that influences who gets an opportunity, including ranking and notifications. For each, record what happens next and how often a human changes it. Then ask, for each, whether you could report the outcome distribution and whether you could explain an individual case from six months ago. Then check the data provenance and the correction path. Then, and only then, look at the models themselves.

Doing it in that order means you discover the scoping gaps, which are the big ones, before spending time on the model that everybody was already thinking about.

If you want this run against your product, get in touch.

Best for

  • Proptech products that have added scoring, ranking, or automation since launch
  • Teams preparing for enterprise procurement, diligence, or an inbound regulatory question

Avoid if

  • The product genuinely makes no decision influencing who gets an opportunity, which is worth confirming rather than assuming

Check before you decide

  • For each output, ask what happens next and how often a human changes it
  • Ask for last quarter's outcome distribution across groups and see how long it takes
  • Ask who has reviewed the ranking or feed system from a fair housing perspective
  • Ask how a decision from six months ago would be explained, and watch the method

Common questions

What is the most common compliance gap in proptech products?

A feature that crossed from informing a decision to making one without anything in the organisation noticing the change. It shows up as a score connected to an automatic action, an advisory recommendation followed in nearly every case, or a customer-configured auto-decline threshold in a settings page nobody has looked at in a year.

Does a human review step always mean a human is deciding?

No. A review that overturns the model in a negligible fraction of cases is not functioning as a review regardless of what the process diagram says, which is why the useful audit question is how often a human actually changes the output rather than whether a review step exists.

What does it mean if a product cannot report its outcome distribution?

That nobody, including the team, knows whether the system produces disparate outcomes, so any assurance offered about fairness is not based on evidence. It also means the work will eventually happen under far worse conditions, when the question arrives from outside, and retrospective fixes only produce data from the point of the fix forward.

Why are ranking and recommendation systems usually excluded from review?

Because the compliance conversation gets framed around screening or valuation, and ranking is filed as a product feature rather than as a system deciding who learns about which housing opportunities. In most organisations nobody has ever reviewed it from a fair housing perspective, and the ranking team has never been in a meeting where the question arose.

In what order should a proptech compliance review run?

List every output influencing who gets an opportunity, including ranking and notifications, and record what happens next for each. Then check whether outcome distribution can be reported and whether an individual decision from six months ago can be explained, then check data provenance and the correction path, and only then examine the models. That order finds the scoping gaps, which are the large ones, before time goes into the model everybody was already thinking about.

Why does third-party data with no provenance create compliance risk?

Because without a record of where each data item came from and when, a dispute can only be handled for the one person who complained. The same erroneous item, whether it is a mismatched record or an expunged item that still appears, continues to affect every other decision that used it, since there is no way to identify which other decisions drew on the same bad data.

Is using a vendor's proprietary scoring model a compliance gap by itself?

Not automatically, but it becomes one when a decision is challenged and the only answer is that a vendor's model produced a number. Whether that is adequate depends on the applicable framework and is a question for counsel, but it is a structurally weaker position than being able to explain the decision, and it should be a considered choice rather than something discovered at the first challenge.

How much work is it to fix a proptech product that cannot report its outcome distribution?

It is real work, because the usual cause is a data model decision made early: outcomes were logged for operations rather than analysis, and group information was never collected in a joinable form. A retrofit only produces usable data from the point of the fix forward, which is why finding this gap early costs far less than finding it after the fact.

Start a project