Building compliant software in real estate / Start here
Why removing the protected field is not a defence
The most common thing a proptech team says when asked about fairness is that the model does not use protected characteristics. It is true, it is well intentioned, and it does not answer the question, because a model does not need a protected attribute to reproduce a protected pattern.
Published August 22, 2026. Editorial.
Key takeaways
- Correlated features act as proxies. Postal code, school district, commute distance, and name-derived signals all carry information about protected characteristics.
- Removing the attribute can make the problem harder to detect, since you lose the ability to measure outcomes across groups.
- The defensible position is measurement of outcomes, not absence of inputs.
- Measuring requires group information, which creates a real tension worth resolving deliberately rather than by avoidance.
"We don't collect race" is offered as a fairness answer more often than any other sentence, and it reflects a genuine and reasonable instinct. It is also, on its own, not a defence, and the reason is worth understanding properly rather than accepting as a slogan.
How a proxy forms without anyone building one
A model learns whatever patterns are present in its training data that help it predict the target. It has no concept of which patterns are acceptable.
Consider a tenant screening model trained on historical outcomes. Among its features is the applicant's current postal code. Residential patterns in many places are strongly correlated with race, for historical reasons that include explicit past discrimination in housing. The model does not know this. It knows that applicants from certain postal codes had different outcomes in the training data, and it uses that.
The result is a model that has never seen a protected characteristic and reproduces a pattern aligned with one. Nobody built a proxy. The pattern already existed in society, and the model found it because finding patterns is the only thing it does.
The same mechanism operates through many features. School district. Distance from the property to the applicant's current address. Employer. The name itself, if any text field carrying it reaches the model. Referral source. Time of application. Device type. Some of these correlations are strong and some are weak, and a model combining several weak signals can reconstruct a protected characteristic effectively even when no single feature would.
This is not a failure of care by the team. It is the default behaviour of statistical learning applied to a world with existing patterns in it.
A worked example, to make the mechanism concrete
Take a tenant screening model whose target is a binary outcome: did the tenant complete the lease term without a serious payment default. The team has, sensibly, excluded race, sex, and familial status from the training data.
The feature set includes prior address history. Among the features derived from that history is average distance between consecutive addresses, on the theory that frequent long-distance moves might correlate with instability. It is a plausible engineering idea and nobody flags it as sensitive, because distance between addresses says nothing explicit about anyone's protected characteristics.
But households with children move for reasons unrelated to financial instability: a larger unit, a specific school catchment, proximity to a co-parent after a separation. If those moves cluster at particular distances, and if the training population's default rates happen to correlate with that pattern for reasons that have nothing to do with the feature itself, the model will pick up a relationship between move distance and default risk that is, underneath, a relationship between familial status and something else entirely. Nobody built a familial-status proxy. The feature was chosen for a reasonable-sounding reason, and the correlation arrived on its own, from the interaction between an ordinary feature and the composition of the training population.
This is why a fairness review that only reads the list of feature names and checks that none of them is named after a protected characteristic will not find this. The feature name is "average move distance." Nothing about that name suggests a problem, and a reviewer working from names alone would approve it. Finding the actual relationship requires measuring the outcome, which is the argument for the rest of this page.
Why removing the attribute can make things worse
There is a further problem with the instinct to delete the field, and it is counterintuitive enough to be worth stating directly.
If you hold no information about the protected characteristic at all, you cannot measure whether your outcomes differ across groups. You have removed your ability to detect the problem while leaving the problem in place. The model still produces whatever pattern it produces, and now nobody can see it.
That is a materially worse position than knowing. The organisation that measures and finds a disparity can investigate it, understand the driver, and decide what to do. The organisation that measures nothing has the same disparity, no knowledge of it, and no ability to answer the question when it is eventually asked from outside.
The tension this creates, and how to manage it
Measuring outcomes across protected classes requires having some information about protected class membership, and collecting that information raises its own concerns, including from applicants who reasonably do not want to provide it and from privacy obligations that may constrain it.
This tension is real and it does not have a simple engineering resolution. What it has are approaches, and choosing among them is a decision to make with counsel rather than in a daily team meeting.
Some organisations collect the data voluntarily and separately, with clear explanation, keeping it isolated from the model's inputs and available only to the measurement process. Some use aggregate or geographic estimation methods to assess outcomes at a population level without attributing characteristics to individuals. Some rely on periodic external assessment rather than continuous internal measurement.
Each has trade-offs, and each is more defensible than not measuring. The engineering requirement, whichever path is chosen, is that the measurement data is architecturally separated from the model's feature set, so that the thing used to check the model cannot leak into the thing being checked.
That separation is worth enforcing with an assertion rather than a convention, because the natural pressure over time is towards using every available field to improve accuracy.
A second angle: proxies can be introduced after launch, not only at training time
Teams sometimes treat this as a one-time review done before a model ships, which misses a real source of new risk: a proxy relationship can appear after launch even when the model itself has not changed.
Suppose a screening model was trained and reviewed, and no strong proxy relationship was found at the time. Six months later, the product adds a new data source, perhaps an eviction records vendor or a supplemental credit file, and pipes several of its fields into the same model as additional features without retraining from scratch, through an online learning update or a scheduled retrain on fresh data. The new fields were not part of the original review. If one of them correlates with a protected characteristic in ways the original feature set did not, the model can develop a proxy relationship it did not have at launch, and nobody re-ran the check because nobody thought of the retrain as a new model requiring one.
The practical response is to treat any change to the feature set, not just a change to the model architecture or target, as an event that triggers the same proxy check as the original build. A new data source, a new derived feature, or a vendor changing what fields it supplies are all changes to what the model can learn from, and each is a point where a new correlation can enter.
What to do instead of deleting
Four practices, in order of how much they help.
Measure the outcome distribution. This is the most important practice. What proportion of applicants in each group receive each outcome, and how does that change over time? Covered on testing a model for disparate impact.
Examine the features for known proxies. Postal code is the obvious one, and if it is in your feature set, somebody should be able to explain why it is necessary and what would be lost by removing it. Necessary is a real category and this is not an argument for removing features without thinking, but it should be a considered decision with a written reason.
Test the model's ability to predict the protected characteristic. If you can train a simple model on your feature set that predicts group membership well, your main model has access to the same signal. This is a useful diagnostic that few teams run and it is not difficult.
Keep the explanation available. If a decision affects someone's housing, they may ask why, and being able to answer specifically rather than generally is both good practice and, in some contexts, an obligation. Covered on explainability and adverse action.
A note on how to talk about this
One thing worth saying to product and marketing teams: "our model is unbiased" is a claim nobody can support, and making it creates exposure rather than reducing it.
What can be supported is a description of practice. We measure outcomes across groups on this cadence, using this method, against these thresholds, and here is what we do when a threshold is crossed. That is a stronger position in every respect, including the legal one, because it is true and demonstrable, and because it does not assert a property that the next change in the data could prove false.
If you want help setting up the measurement, get in touch.
Best for
- Any product where a model influences who gets housing, credit, or a viewing
- Teams whose current fairness answer is that they do not collect protected attributes
Avoid if
- The measurement approach has not been settled with counsel, since how to obtain group information is a legal question before it is a technical one
Check before you decide
- Ask what proportion of each group receives each outcome, and whether anyone can produce the number today
- Check whether postal code or a close correlate is in the feature set, and whether its inclusion has a written justification
- Train a simple model to predict group membership from your feature set and see how well it does
- Check that measurement data is architecturally separated from model inputs, and asserted rather than assumed
Common questions
Can a model discriminate without using protected attributes?
Yes, because correlated features act as proxies. Residential patterns are strongly correlated with race in many places for historical reasons, so a model using postal code can reproduce a protected pattern despite never seeing a protected characteristic, and it does so because finding predictive patterns is the only thing it does.
Why can removing protected attributes make the situation worse?
Because holding no group information removes your ability to detect a disparity while leaving the disparity in place. An organisation that measures and finds a problem can investigate and act, whereas one that measures nothing has the same problem, no knowledge of it, and no answer when the question comes from outside.
How do you measure outcomes without collecting sensitive data on everyone?
The common approaches are voluntary separate collection with clear explanation and architectural isolation from model inputs, aggregate or geographic estimation that assesses outcomes at population level without attributing characteristics to individuals, or periodic external assessment. Each has trade-offs and each is more defensible than not measuring, and choosing among them is a decision for counsel.
How can you tell whether your features encode a protected characteristic?
Train a simple model on your feature set to predict group membership and see how well it performs. If a basic model predicts it accurately, your main model has access to the same signal, and this diagnostic is straightforward to run yet rarely done.
Should a product claim its model is unbiased?
No, because it is a claim nobody can support and asserting it creates exposure rather than reducing it. Describing the practice instead, including the measurement cadence, method, thresholds, and the response when a threshold is crossed, is stronger in every respect because it is true, demonstrable, and not proven false by the next change in the data.
What is a proxy feature in a housing model?
A feature that carries information about a protected characteristic without naming it directly, such as postal code, school district, commute distance, or a name-derived signal. A model learns whatever patterns predict its target, so if residential patterns correlate with race for historical reasons, the model uses that correlation without knowing what it represents.
Deleting a protected field versus measuring outcomes: which is the better defence?
Measuring outcomes, because deleting the field only removes your ability to detect a disparity while the underlying pattern stays in the model. An organisation that measures and finds a disparity can investigate and respond. One that deletes the field and measures nothing has the same disparity with no way to know about it or answer for it later.
How do you test whether a model's features encode a protected characteristic?
Train a simple model on the same feature set to predict group membership directly. If that simple model performs well, the main model has access to the same signal through its features, even with the protected attribute itself absent. This diagnostic is straightforward to run and is skipped by most teams building housing models.
Related reading
How to tell if an AI feature idea is worth building
Most AI feature ideas look good in a demo and fail during the work of making them reliable. Here are the four questions we ask to tell the ones worth building from the ones that just look good in a demo.
Never let the model grade its own work
When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.
AI is an amplifier, not a fix
The 2025 DORA data says AI raises throughput and hurts stability at the same time. Which of those two you get is decided by what has to pass before a change is released, and by nothing else.