Building compliant software in real estate / The rules that apply to your code

The rules that apply to your code

Tenant screening and fair housing

Screening software is used at the point where a person does or does not get housing. HUD's May 2024 guidance addressed how the Fair Housing Act applies to that function, including where algorithms and AI perform it, and the design implications are concrete.

Published August 22, 2026. Editorial.

Key takeaways

  • HUD's guidance addresses fair housing issues in screening practices including the use of third-party screening companies and machine learning.
  • The guidance describes screening companies as helping to implement rather than effectively set a housing provider's policies, and it favours customisable criteria.
  • Criteria that are transparent and stated in plain language are easier to defend than a score whose basis nobody can articulate.
  • A single criterion applied uniformly to everyone can still produce a disparate impact, which is the point of measuring outcomes rather than reviewing rules.

In May 2024, HUD issued guidance addressing the application of the Fair Housing Act to tenant screening, including the increasing use of third-party screening companies and the emerging use of machine learning and artificial intelligence in that function [1].

For a team building screening software, several themes in that guidance translate into design decisions worth making deliberately.

Who sets the policy

A theme in the guidance is that screening companies should serve to help implement, rather than effectively set, a housing provider's screening policies, and should offer customisability as to criteria, standards, and weights [1].

That is a product architecture statement as much as a compliance one, and it goes against how screening products are usually built. The commercially attractive design is a single proprietary score: it is easier to build, easier to sell, and it makes the vendor's model the differentiator. It also has the effect of the vendor setting the criteria for every provider using it.

The design the guidance favours is different. Criteria are configurable by the provider. Weights are visible and adjustable. The provider can see what the criteria are and take responsibility for them. The product supports the provider's policy rather than substituting for it.

There is a real tension here for a vendor, because configurability means providers can configure badly. The reasonable compromise is to make the configuration explicit and reviewable, provide defaults with stated reasoning, show the likely impact of a configuration before it is applied, and keep a record of what each provider selected and when.

Transparency of criteria

The guidance notes that criteria used for determining eligibility need to be transparent and use plain writing [1].

For a build, that means the criteria have to exist in an articulable form. A rule engine with named rules satisfies this naturally. A learned model producing a score does not, unless substantial work goes into making its basis explainable, which is covered on explainability and adverse action.

This is a genuine architectural choice and it is worth facing early rather than discovering after building. A transparent rule system is less powerful and far easier to defend. A learned model may be more accurate and requires an explanation layer, ongoing measurement, and a much stronger governance process. Neither is automatically right, and the choice should be made with the compliance implications visible rather than as a purely technical preference.

Our general view, offered as engineering judgment rather than legal advice, is that in decisions this consequential the more complex option should have to prove it is better. If a learned model is not measurably better than a transparent rule set on the outcomes that matter, the transparent one is the better product.

Uniform criteria can still produce disparate outcomes

The most important conceptual point, and the one that surprises engineering teams most.

A criterion applied identically to everyone can still produce substantially different outcomes across protected groups, because the underlying populations differ in ways shaped by history and circumstance. A rule that looks neutral at first sight is not necessarily neutral in effect.

This is why reviewing the rules is insufficient, and why measuring outcomes is the thing that actually answers the question. A screening product that cannot report the distribution of outcomes across groups cannot tell whether its neutral-looking criteria produce a neutral-looking result, and neither can its customers.

A concrete example of a criterion that looks neutral and is not necessarily neutral in effect: a minimum income-to-rent ratio, applied at the same number to every applicant. The criterion does not mention any protected characteristic and treats every application the same way arithmetically. But if the applicant pool includes households relying on a housing voucher or another form of assistance, and if the criterion is calculated only against wage income without accounting for that assistance, the practical effect can fall unevenly across household types in ways that track familial status or other protected characteristics, depending on who in the local population relies on that kind of support. Whether a specific formulation of an income criterion is lawful in a specific jurisdiction is a legal question, not an engineering one, and the point here is narrower: a rule can be applied with perfect consistency and still need its outcomes checked, because consistency of application is not the same thing as neutrality of effect.

The mechanics of that measurement are on testing a model for disparate impact. The design requirement is that the product can produce the report at all, per provider and in aggregate, which is a data model decision made early.

Data quality is a fairness issue

Screening decisions depend on data from third parties, and that data has error rates.

A record matched to the wrong person, an expunged item that still appears, a satisfied judgment shown as outstanding, an address history with gaps: each of these produces a wrong decision about a real person's housing. And matching error rates are not uniformly distributed, which means data quality problems can themselves produce disparate outcomes even when the criteria are sound.

Three engineering practices help. Record the provenance and retrieval date of every data element used in a decision, so a disputed item can be traced. Make the matching logic conservative and record its confidence, rather than accepting a probable match silently. And build a correction process that reaches every affected decision: when a data item is disputed and corrected, the affected decisions should be identifiable, which requires having recorded which decisions used which data.

That last capability is rarely built and is what makes a dispute process real rather than nominal.

Name matching is worth calling out specifically, because it is a common source of the wrong-person problem and it interacts with fairness in a way that is easy to miss. Common surnames, generational naming patterns, and transliteration from other naming conventions all increase the rate at which one applicant's record is confused with another's in a database match. If those naming patterns are not evenly distributed across the population, and they generally are not, the wrong-match rate itself is not evenly distributed either, which means a purely technical weakness in matching logic can produce a pattern that looks like the model discriminating even when no criterion or feature does. Conservative matching, meaning a match is accepted only above a confidence threshold and anything below it routes to manual verification rather than an automatic accept or reject, reduces this specific failure mode directly.

Human review, where it matters

The strongest design property for a screening product is that adverse outcomes can be reviewed by a person with enough context to reverse them.

That means the interface presents not just the outcome but the specific factors that caused it, the underlying data with its source, and a clear path to override with a recorded reason. It also means the product should make it easy for a provider to route certain categories to review rather than auto-declining, and should probably make that the default for the outcomes with the most severe consequences.

Auto-decline at scale is the configuration that turns a decision support tool into a decision maker, as described on what regulation reaches a real estate product. It is a legitimate feature that customers want, and it deserves a deliberate design conversation rather than arriving as a checkbox.

A design pattern worth considering directly, rather than as an afterthought once auto-decline already exists as a feature: tiered thresholds rather than a single cutoff. An application well clear of every criterion can be approved automatically with little risk of a wrong outcome. An application that fails a single criterion by a narrow margin, or that fails only because of a data item with a low match confidence, is a different case, and routing it to a person rather than to an automatic decline costs the provider a small amount of review time in exchange for a lower risk of a wrong and consequential outcome. Building the threshold logic to support this distinction from the start, rather than only a binary pass or fail, is a modest amount of engineering work that changes where that risk falls.

It is also worth designing the override interface so that a reviewer overturning an automatic decision is not fighting the tool to do it. If reversing an automatic decline requires more steps than accepting it, reviewers will accept it more often regardless of whether that is the right outcome for the applicant in front of them, simply because the path of least resistance shapes behaviour under time pressure. An interface where override is at least as easy as acceptance removes that bias from the design itself rather than relying on the reviewer to fight it every time.

Fair housing determinations are fact-specific and this page is engineering guidance rather than legal advice. If you are building screening software and want the measurement and explanation layers designed properly, get in touch.

Best for

  • Screening products where the provider should own the criteria rather than inheriting a vendor's
  • Teams choosing between a transparent rule engine and a learned scoring model

Avoid if

  • The product cannot currently report outcome distributions, in which case that capability comes before any further model work

Check before you decide

  • Ask whether the provider can see and adjust the criteria, or only the threshold
  • Ask whether the product can report outcome distribution per provider
  • Check whether a disputed and corrected data item can be traced to the decisions that used it
  • Check what happens on an adverse outcome: what the reviewer sees and whether override is recorded

Common questions

Does the Fair Housing Act apply to algorithmic tenant screening?

HUD issued guidance in May 2024 addressing the Act's application to tenant screening including the use of third-party screening companies and the emerging use of machine learning and artificial intelligence. Performing the function with a model does not move the obligation elsewhere.

Should a screening product use a single proprietary score?

HUD's guidance favours screening companies helping to implement rather than effectively set a provider's policies, and favours customisability as to criteria, standards, and weights. A single opaque score has the effect of the vendor setting criteria for every provider using it, which goes against that guidance even though it is the commercially easier product to build and sell.

Can criteria applied equally to everyone still be a problem?

Yes, and this is the point engineering teams most often miss. A criterion applied identically can produce substantially different outcomes across groups because the underlying populations differ in ways shaped by history, which means reviewing the rules cannot answer the question and measuring outcomes is what does.

Why is third-party data quality a fairness issue in screening?

Because errors like a record matched to the wrong person, an expunged item still appearing, or a satisfied judgment shown as outstanding each produce a wrong decision about someone's housing, and matching error rates are not uniformly distributed across populations. That means data quality problems can produce disparate outcomes even when the criteria themselves are sound.

What makes a dispute process real rather than nominal?

The ability to identify which decisions used which data, so that when an item is disputed and corrected the affected decisions can be found and revisited. This requires recording provenance and retrieval dates at decision time, and it is rarely built, which is why most correction paths fix the record and leave the decisions it produced untouched.

What is the difference between a rule-based screening engine and a learned scoring model?

A rule engine with named rules states its criteria in an articulable form that satisfies HUD's guidance on plain-language transparency directly. A learned model producing a score does not, unless substantial work goes into making its basis explainable. A rule engine is less powerful but far easier to defend, while a learned model may be more accurate and requires an explanation layer and stronger governance.

How should auto-decline thresholds be configured in a screening product?

Auto-decline at scale is the configuration that turns a decision support tool into a decision maker, so it deserves a deliberate design conversation rather than arriving as a default checkbox. A reasonable approach makes human review the default for the outcomes with the most severe consequences and routes only lower-stakes categories to automatic action.

Is it worth using a proprietary scoring vendor instead of building screening logic in house?

The answer depends on whether the vendor's criteria are configurable and explainable, not just accurate. HUD's guidance favours screening companies helping implement rather than effectively setting a provider's policies, so a single opaque vendor score that the provider cannot adjust or explain goes against that guidance, even though it is the easier product to buy.

Start a project