Building compliant software in real estate / Building it right
Testing a model for disparate impact
Fairness measurement in housing products usually happens once, in a review, producing a document. The obligation is better served by measurement that runs continuously, like any other check, because a model's behaviour changes as its data changes and a good result in March says little about September.
Published August 22, 2026. Editorial.
Key takeaways
- Which fairness metric applies is a legal and policy decision, not an engineering one. Different definitions conflict mathematically.
- Measure outcomes across groups continuously, with the results retained long enough to show a trend.
- Segment the measurement, because an aggregate result conceals a problem confined to one region, price band, or property type.
- The measurement pipeline must be architecturally separated from the model's inputs, and that separation asserted rather than assumed.
This page is about the mechanics of measurement. It deliberately does not tell you what fair means, because that is not an engineering question and treating it as one is how teams end up with a defensible-looking number that answers the wrong thing.
The part that is not yours to decide
There are several mathematical definitions of fairness, and they are not compatible with each other. Equal outcome rates across groups, equal error rates across groups, and equal predictive value across groups are different properties, and results in the fairness literature establish that you generally cannot satisfy all of them simultaneously when base rates differ across groups.
That is not a technicality. It means somebody has to choose, the choice has consequences for real people, and the appropriate choice depends on the specific context, the legal framework, and what the decision is used for. It should be made with counsel and documented with reasoning, in the same format as the decision records described on the healthcare hub at Required vs Addressable safeguards.
What goes wrong when engineering decides is subtle and common: a library is chosen, its default metric is adopted because it is the default, and the organisation now has a fairness programme optimising a property nobody selected on purpose.
A worked illustration of why the choice matters: suppose a screening model approves loans or leases at a lower rate for one group than another, but among the applicants it approves, the outcomes are equally good across both groups. A metric based on equal approval rates would flag this as a problem. A metric based on equal predictive accuracy, sometimes called calibration, would not, because the model is equally right about the people it approves in both groups; it is simply more conservative about one group at the point of decision. Both descriptions are true of the same model at the same time, and which one matters more depends on what the decision is for and what the legal framework requires, not on which library shipped with which default. An engineering team asked to "make the model fair" without being told which of these properties to target has been handed a policy question presented as a technical one, and no amount of engineering skill answers a question like that.
So: policy chooses the metric and the threshold. Engineering makes the measurement continuous, correct, visible, and hard to disable. That division is the whole structure of this page.
What to measure
Once the metric is chosen, the measurement itself is ordinary data work.
For each decision the model influences, record the outcome, the group information available through whatever method was chosen, and the segment attributes. Then compute the chosen metric across groups, on a schedule, and store the result with a timestamp.
Three details make the difference between a measurement that finds things and one that reassures.
Segment it. An aggregate number across a national product can look fine while a serious disparity exists in one metropolitan area, one price band, or one property type. Compute the metric within segments as well as overall, and set the segment definitions deliberately rather than by whatever dimensions happen to be in the data.
Retain the series. The most valuable signal is the trend, and a single point tells you almost nothing. A metric that has changed steadily over eight months is a much clearer finding than any single measurement, and it is only visible if results are kept.
Include the decisions that did not happen. Applications abandoned partway, users who never saw a listing, requests routed to manual review and never completed. These are outcomes too, and measuring only completed decisions systematically excludes a category where disparities often show up.
A specific version of this worth watching for: an application form long enough, or a required document difficult enough to produce, that some applicants abandon it before a decision is ever recorded. If abandonment happens at different rates across groups, for reasons connected to the form itself rather than to anything about the applicant's likely outcome, a measurement counting only completed applications will look clean while the product is filtering people out earlier in the process where nothing is being measured at all. The fix is to instrument the funnel, not only the decision, recording when and where an applicant stops rather than only what happens to the ones who finish.
Where the measurement data is stored
The data used to check the model must not be reachable by the model.
This sounds obvious and is violated constantly, because the natural pressure in a feature engineering process is to use every available field, and because a well-meaning engineer adding a feature will not know that a particular column exists for measurement purposes only.
The architectural answer is separation: the group information is kept in a store the training and inference pipelines cannot read, joined only in the measurement process. The practical answer is to assert it, with a check that fails if a protected attribute or its designated proxy appears in the model's feature set. People stop following conventions over time; an assertion is checked on every change.
Testing before deployment, not only after
Continuous measurement of production outcomes tells you what happened. A pre-deployment check tells you before it happens to anyone.
The method is a held-out evaluation set with group information attached, run against any candidate model before it is released, computing the same metric that production measures. A model whose fairness metric degrades relative to the current production model is not released, in the same way a model whose accuracy degrades is not released.
Making this a check that blocks the release, rather than a report, is the difference between a fairness programme that constrains and one that observes. And it is worth writing down what happens when the check fails, before it fails, because the conversation is much harder to have well under deadline pressure with a model everybody has already been told about.
Keeping it honest
A fairness check is a control, and controls that the people they constrain can disable without anyone noticing are not controls. The same logic as using evals as compliance evidence applies more strongly here, because the pressure to relax a fairness threshold is real and arrives at the least convenient moment.
Concretely: changes to the metric definition, the threshold, or the segment definitions get reviewed separately from the model change that motivated them. The count of segments excluded from measurement is itself a tracked number that cannot rise without anyone noticing. Failures are recorded even when subsequently overridden, with the override reason. And the person who approves a threshold change is not the person releasing the model that needs it.
None of that is exotic governance. It is the same structure as any other control that matters.
What to do with a finding
Finding a disparity is the beginning of work, not the end, and the response depends entirely on the cause. So the first step after detection is investigation: which feature or features cause the difference, is it present in the training data, is it a data quality artefact, is it concentrated in a segment.
The responses available range from feature changes, to reweighting or resampling training data, to constraining the model, to changing the decision threshold, to routing affected cases to human review, to not deploying. Which is appropriate is a judgment involving legal, product, and engineering together.
What matters from an engineering standpoint is that the finding is recorded, the investigation is recorded, and the decision is recorded, with dates. An organisation that measures, finds something, investigates, and documents its response is in a substantially stronger position than one that never looked, and it is also simply doing the right thing by the people the decision affects.
One more distinction worth drawing before deciding on a response: a disparity traced to the training data reflecting a real, historical pattern in the world is a different problem from a disparity traced to a bug, such as a matching error or a mislabelled outcome in the data pipeline. The second kind should be fixed outright, and fixing it is uncontroversial. The first kind is harder, because the model is behaving exactly as trained and the pattern it learned is genuinely present in history, which means the available responses (reweighting, constraining, threshold changes, routing to human review) are all choices about how much the model's future behaviour should be allowed to continue a historical pattern rather than correct for it. That choice belongs with the people who decided the metric and the threshold in the first place, for the same reason they decided those: it is a policy decision presented as an engineering fix, and treating it as a bug to be patched skips the part where somebody with the authority to make that call actually makes it.
If you want help building this properly, get in touch.
Best for
- Any housing product where a model influences who gets an opportunity
- Teams with a periodic fairness review that they want to turn into a continuous control
Avoid if
- The metric and threshold have not been chosen with counsel, since building the pipeline first tends to fix in place whichever default the library comes with
Check before you decide
- Ask who chose the fairness metric and whether the reasoning is written down
- Check whether results are segmented and whether the series is retained long enough to show a trend
- Check for an assertion that protected attributes cannot appear in the model's feature set
- Ask whether a fairness regression blocks a release or produces a report
Common questions
Which fairness metric should a housing model use?
That is a legal and policy decision rather than an engineering one, because different definitions conflict mathematically and generally cannot be satisfied simultaneously when base rates differ across groups. The common failure is adopting whichever metric a library implements by default, which leaves the organisation optimising a property nobody chose on purpose.
Why does fairness measurement need to be continuous?
Because a model's behaviour changes as its inputs and its population change, so a good result in one quarter says little about the next. The most valuable signal is the trend rather than any single measurement, which means results have to be retained as a series rather than produced as a one-time document.
Why segment fairness measurements?
Because an aggregate number across a national product can look entirely acceptable while a serious disparity exists in one metropolitan area, price band, or property type. Segments should be defined deliberately rather than following whatever dimensions happen to be present in the data.
How do you stop measurement data leaking into the model?
Architecturally, by keeping group information in a store the training and inference pipelines cannot read and joining it only in the measurement process, and then asserting it with a check that fails if a protected attribute or designated proxy appears in the feature set. People stop following conventions because feature engineering naturally tries to use every available field.
What should happen when a fairness check fails before deployment?
The model should not be released, in the same way one with degraded accuracy would not, and what happens next should be written down before the first failure rather than negotiated under deadline pressure. Making it a check that blocks the release, rather than a report, is the difference between a programme that constrains and one that merely observes.
What is the difference between pre-deployment and continuous fairness testing?
Pre-deployment testing runs a candidate model against a held-out evaluation set with group information attached before it is released, so a model whose fairness metric degrades does not go live. Continuous measurement tracks production outcomes after deployment, since a model's behaviour changes as its inputs and population change over time. A complete programme needs both, not one in place of the other.
How should a team respond when a disparate impact finding appears?
Treat it as the start of an investigation, not the end of one. First identify which feature or features cause the difference and whether it traces to training data, a data quality artefact, or one segment. The available responses range from feature changes to reweighting training data to routing affected cases to human review, and the choice involves legal, product, and engineering together.
Is disparate impact testing worth the engineering cost for a small housing product?
The cost grows with product complexity, but the underlying obligation, that outcomes should be measurable and explainable, does not depend on size. A smaller product with fewer decisions can build the same continuous measurement pattern at proportionally lower cost, and starting early avoids the retrofit problem where historical decisions were never recorded with group information attached.
Related reading
Never let the model grade its own work
When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.
How to tell if an AI feature idea is worth building
Most AI feature ideas look good in a demo and fail during the work of making them reliable. Here are the four questions we ask to tell the ones worth building from the ones that just look good in a demo.
Adding eval tests was the best decision we made
We generate every line of code we release. The single change that made that safe was writing the check before the code, and it turned out to change how we review, how we spec, and how fast we can work.