Building compliant software in real estate / The rules that apply to your code
AVM quality control standards in practice
The AVM rule is unusual in that it names testing as a requirement rather than leaving it implicit. For an engineering team, that makes it one of the easier regulations to build for, because most of what it asks for is measurement you can automate.
Published August 22, 2026. Editorial.
Key takeaways
- Five quality control factors: confidence in estimates, protection against data manipulation, avoidance of conflicts of interest, random sample testing and reviews, and compliance with nondiscrimination laws.
- The rule is deliberately flexible about how, which means you owe a written decision about your approach for each factor.
- Random sample testing is a named requirement and matches a continuous evaluation process directly.
- Data manipulation protection is a security and integrity requirement on the inputs, not only on the model.
The interagency final rule implementing section 1473(q) of the Dodd-Frank Act took effect on October 1, 2025 [1]. It applies to mortgage originators and secondary market issuers using automated valuation models to determine the collateral worth of a mortgage secured by a consumer's principal dwelling, in certain credit decisions or securitization determinations.
The rule requires adopting policies, practices, procedures, and control systems to ensure covered AVMs adhere to quality control standards designed to: ensure a high level of confidence in the estimates produced; protect against the manipulation of data; seek to avoid conflicts of interest; require random sample testing and reviews; and comply with applicable nondiscrimination laws [1].
It deliberately does not prescribe how, giving institutions flexibility based on size and the risk and complexity of the transactions involved [1]. That flexibility, as in the HIPAA Security Rule, means you owe a documented decision for each factor rather than an exemption.
The practical consequence of that flexibility is worth spelling out, because it is easy to read as permission to do nothing formal. A small institution using a single third-party AVM at low volume can reasonably choose a lighter approach to random sample testing than a large issuer running its own model across a national portfolio, and the rule allows for that difference. What it does not allow is skipping the decision about what approach fits the institution's own size and risk. The written record that says "given our volume and risk profile, we sample at this rate and review at this cadence" is the artefact the rule is actually asking for, regardless of how light or heavy that approach turns out to be.
Here is what each factor asks of a build.
Confidence in the estimates
The first factor is about the model being good at its job in a way you can demonstrate.
What this looks like in engineering terms is straightforward: a held-out evaluation set, accuracy metrics tracked over time, and confidence intervals rather than point estimates. The part teams build too little of is the segmentation. A model with strong overall accuracy can be substantially worse on particular property types, price bands, geographies, or transaction contexts, and an aggregate metric conceals that entirely.
So the check is not one accuracy number. It is accuracy by segment, with the segments chosen to include the ones where you have less data, because sparse segments are where confidence degrades and where aggregate reporting hides it.
A concrete case worth walking through: a valuation model trained mostly on suburban single-family homes, because that is where most transactions in the training data came from, will typically show strong aggregate accuracy while performing worse on rural properties, condominiums with unusual fee structures, or newly built homes with no comparable sales history yet. If the reporting stops at the aggregate number, none of that shows up, and the model keeps producing confident-looking estimates for exactly the property types it understands least. The fix is not a different model. It is reporting that makes the weak segments visible instead of averaging them away.
The second part is knowing when the model should decline to answer. A valuation model that produces a number for every input, including inputs unlike anything in its training data, is producing false confidence. Building an explicit abstention path, where low-confidence cases route to human appraisal rather than returning a number, is both better engineering and a much easier thing to defend.
Protection against manipulation of data
This factor is often read as a model concern and it is mostly an input concern.
The question is whether someone with an interest in the outcome can influence the inputs the model uses. That covers a range of concrete mechanisms: whether property characteristics can be edited by a party to the transaction, whether comparable sales data comes from a source that can be filled with false entries, whether user-supplied information about condition or improvements feeds the estimate, and whether the same person can both request a valuation and modify the data it draws on.
The engineering answers are ordinary integrity controls. Provenance tracking on every input, so you can say where each figure came from. Immutable records of what the model saw at the time it produced a given estimate, which also serves the auditability requirement. Separation of duties between requesting a valuation and editing underlying data. And monitoring for the pattern where an input is edited shortly before a valuation is requested, which is the typical sign of the behaviour this factor exists to catch.
That last one is worth building as an alert rather than a report, because it is actionable in the moment and useless six months later.
Avoiding conflicts of interest
Largely a process and organisational factor rather than an engineering one, but it has a technical component worth naming.
If the system allows the party who benefits from a higher valuation to influence the valuation process, whether by selecting the model, choosing the comparables, re-running until a satisfactory number appears, or adjusting inputs, that is a design problem. The engineering control is to record every valuation request and its result, including the ones that were discarded, so that a pattern of re-running until the number is right is visible rather than invisible.
Systems that only store the accepted valuation cannot show this, and the absence is difficult to distinguish from good behaviour.
A related pattern worth checking for specifically: a workflow where a valuation request can be cancelled and immediately resubmitted with different inputs, such as a corrected square footage or an updated list of recent improvements, without any link recorded between the original request and the resubmission. Each individual resubmission may be entirely legitimate, since inputs genuinely do need correction sometimes. What makes it a conflict of interest risk is the absence of a link: without one, a party who resubmits five times and keeps only the highest result looks, in the data, identical to five unrelated homeowners who each requested a valuation once. Recording a session or request-chain identifier that survives a cancel-and-resubmit cycle is a small schema change that makes the pattern visible.
Random sample testing and reviews
This is the factor that matches most directly how we build, because it describes a continuous evaluation process in regulatory language.
A defensible implementation samples valuations on an ongoing basis, compares them against an independent standard such as a subsequent appraisal or an actual sale price where available, records the comparison, and tracks the distribution of error over time. Random selection matters, because a sample chosen by anyone with an interest in the result is not a sample.
Two things make this substantially more useful. Stratify the sample so that sparse segments are represented, since pure random sampling from a skewed population will under-sample exactly the cases where the model is weakest. And retain the results for the long term, because the value is in the trend, and a gradual change in the error distribution over a year is the signal you most want and the one short retention windows delete.
Compliance with applicable nondiscrimination laws
The fifth factor is a quality control standard requiring covered AVMs to comply with applicable nondiscrimination laws [1], and it is the one that cannot be satisfied by anything except measurement.
The mechanics are on testing a model for disparate impact, and the reason removing protected attributes does not address it is on why removing the protected field is not a defence.
The specific concern in valuation has a long documented history in housing, and it is why this factor was included. A model trained on historical valuations learns historical patterns, including any that reflect past discrimination in the market it learned from. The model is not doing anything wrong in a technical sense; it is accurately reproducing what was in the data. That is precisely the problem.
Making it a build artefact rather than a review
The pattern we would recommend, and it is the same one across every regulated industry: turn each factor into checks that run continuously, keep a written decision record for the approach you chose to each factor, and retain the results for the long term.
Doing that produces the policies, practices, procedures, and control systems the rule asks for, as a byproduct of building, rather than as a document written alongside a system that may or may not implement it. The reasoning is on using evals as compliance evidence, and it applies with unusual directness here because this rule explicitly names testing.
Whether the rule applies to your institution and which of your models are covered are questions for counsel. If you want help building the measurement once that is settled, get in touch.
Best for
- Valuation models used in credit or securitization workflows at covered institutions
- Teams that have accuracy metrics but no segmented reporting or sampling process
Avoid if
- Whether your models are covered has not been determined, since scope decides everything that follows
Check before you decide
- Ask for accuracy by segment rather than in aggregate, including the sparse segments
- Ask whether the model can decline to answer, and what happens when it does
- Check whether discarded valuation runs are recorded or only accepted ones
- Check whether sample test results are retained long enough to show a trend
Common questions
What are the AVM quality control factors?
Five: ensuring a high level of confidence in the estimates produced, protecting against the manipulation of data, seeking to avoid conflicts of interest, requiring random sample testing and reviews, and complying with applicable nondiscrimination laws. The rule took effect on October 1, 2025 and deliberately does not prescribe how each is achieved.
How do you demonstrate confidence in a valuation model's estimates?
With accuracy tracked by segment rather than in aggregate, since a model with strong overall numbers can be substantially worse on particular property types, price bands, or geographies and an aggregate metric conceals that. An explicit abstention path, where low-confidence cases route to human appraisal instead of returning a number, is both better engineering and easier to defend.
What does protecting against data manipulation mean technically?
It is mostly an input concern rather than a model concern: whether a party with an interest in the outcome can influence what the model sees. The controls are provenance tracking on every input, immutable records of what the model saw when it produced a given estimate, separation between requesting a valuation and editing underlying data, and alerting on inputs edited shortly before a valuation request.
How should random sample testing be implemented?
Sample valuations on an ongoing basis, compare against an independent standard such as a later appraisal or actual sale price, record the comparison, and track the error distribution over time. Stratify so sparse segments are represented, since pure random sampling under-samples exactly the cases where the model is weakest, and retain results long term because the value is in the trend.
Why does the AVM rule include a nondiscrimination factor?
Because a model trained on historical valuations learns historical patterns, including any reflecting past discrimination in the market it learned from. The model is not malfunctioning when it does this, it is accurately reproducing its training data, which is exactly why the factor cannot be satisfied by anything other than measuring outcomes.
Which institutions does the AVM rule apply to?
Mortgage originators and secondary market issuers that use automated valuation models to determine the collateral worth of a mortgage secured by a consumer's principal dwelling, in certain credit decisions or securitization determinations. It was adopted by the OCC, Federal Reserve, FDIC, NCUA, CFPB, and FHFA jointly and took effect on October 1, 2025.
How is preventing data manipulation different from ensuring valuation accuracy?
Accuracy is about whether the model's estimates are good, checked through held-out evaluation and segmented metrics. Manipulation protection is about whether a party with an interest in the outcome can influence the inputs the model sees, checked through provenance tracking, immutable input records, and separation between requesting a valuation and editing the data behind it.
What does it cost to build the AVM rule's random sample testing requirement?
The rule does not specify a cost or a fixed process, deliberately leaving flexibility based on an institution's size and the risk and complexity of its transactions. In engineering terms, the requirement matches a continuous evaluation pipeline that samples valuations, compares them against an independent standard, and retains the error distribution over time, which is ordinary measurement infrastructure rather than a one-time project.
Related reading
How to tell if an AI feature idea is worth building
Most AI feature ideas look good in a demo and fail during the work of making them reliable. Here are the four questions we ask to tell the ones worth building from the ones that just look good in a demo.
Adding eval tests was the best decision we made
We generate every line of code we release. The single change that made that safe was writing the check before the code, and it turned out to change how we review, how we spec, and how fast we can work.
More in The rules that apply to your code
Tenant screening and fair housing
Screening software is used at the point where a person does or does not get housing. HUD's May 2024 guidance addressed how the Fair Housing Act applies to that function, including where algorithms and AI perform it, and the design implications are concrete.
Housing advertising and audience targeting
This is the obligation engineers find least intuitive, because audience targeting and delivery optimisation feel like technical concerns. They are not. Deciding who sees a housing listing is deciding who learns the opportunity exists, and that is advertising of a housing opportunity.