Budget and rules

What regulators and standards bodies say about testing AI

As of 30 September 2026, the EU AI Act states a testing duty for AI directly: Article 9 requires high-risk AI systems to be tested before they are placed on the market, against metrics and thresholds defined in advance. After an amendment published on 24 July 2026, that duty applies from 2 December 2027 or 2 August 2028, depending on the type of system. In the United States, NIST's AI Risk Management Framework describes testing before release and during operation, and it is voluntary. ISO/IEC 42001 is a management standard, and Colorado's 2026 law asks for documentation and records. This page describes what the texts say. It is general information, and a lawyer should confirm what applies to your product.

Published September 30, 2026. Editorial.

Key takeaways

  • EU AI Act Article 9(8) says high-risk AI systems are tested before they are placed on the market, against metrics and thresholds defined in advance. The act sets no numeric accuracy level.
  • Regulation (EU) 2026/1744, published 24 July 2026, moved the high-risk rules, including Articles 9 and 15, from 2 August 2026 to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems.
  • NIST's AI Risk Management Framework 1.0 is voluntary. It says AI systems should be tested before deployment and regularly while in operation, and NIST says, as of 30 September 2026, that version 1.0 is being revised.
  • Colorado's SB24-205 never applied in its 2024 form. SB26-189 replaced it with duties from 1 January 2027, and the legislature's summary of that act states no testing duty.
  • A dated eval run, with its test set, threshold and result, is the type of record these texts describe. Whether it meets a specific legal duty is a question for a lawyer.

On 24 July 2026 the European Union published a regulation that changed when its AI testing rules apply. Any document that was written before that date and still gives 2 August 2026 for the EU AI Act's high-risk rules is out of date. Colorado's AI law has been changed twice since it was signed. For that reason every fact on this page is given with the date it was read: 30 September 2026.

This page describes what the texts say about testing and evaluation. It is general information. For legal advice on your own product, ask a lawyer, who should confirm which of these texts apply and what they require. The EU act and the NIST documents say "testing", "evaluation" or "measurement" where the industry says "eval", its short name for a repeatable test of an AI system. That term is explained in what an AI eval is.

The instruments in one table

Every row is as of 30 September 2026.

Instrument Issued by What it says about testing When it applies
EU AI Act, Articles 9 and 15 European Parliament and Council High-risk systems are tested before release, against thresholds defined in advance 2 December 2027 or 2 August 2028, by type of system
EU AI Act, Article 55 European Parliament and Council Providers of the largest general-purpose models perform model evaluation 2 August 2025
General-Purpose AI Code of Practice Independent experts; described on a European Commission page A voluntary way for model providers to show compliance Voluntary; published 10 July 2025
AI Risk Management Framework 1.0 NIST (United States) Test before deployment and regularly in operation; document test sets and metrics Voluntary; released 26 January 2023; being revised
Generative AI Profile, AI 600-1 NIST (United States) Set minimum thresholds and review them when approving deployment Voluntary; released 26 July 2024
ISO/IEC 42001:2023 ISO and IEC A management system standard; text not read for this page Voluntary; published 18 December 2023
Colorado SB26-189 Colorado General Assembly The summary states documentation and records, and no testing duty 1 January 2027

What the EU AI Act says about testing

Regulation (EU) 2024/1689, the Artificial Intelligence Act, was published on 12 July 2024 [1]. Its testing rules are in Chapter III, Section 2, and they cover one legal category: "high-risk AI systems", defined in Article 6 and in Annexes I and III of the act. Most business software is outside that category. Our page on the EU AI Act and your product explains the risk classes. For the dates, use the next section of this page.

Article 9 has three passages about testing [1].

  • Article 9(6): "High-risk AI systems shall be tested for the purpose of identifying the most appropriate and targeted risk management measures."
  • Article 9(8), on timing: testing "shall be performed, as appropriate, at any time throughout the development process, and, in any event, prior to their being placed on the market or put into service."
  • Article 9(8), on method: "Testing shall be carried out against prior defined metrics and probabilistic thresholds that are appropriate to the intended purpose of the high-risk AI system."

Article 9(2) describes the risk management system that this testing belongs to as "a continuous iterative process planned and run throughout the entire lifecycle of a high-risk AI system" [1].

Article 15 is about results. Article 15(1) says high-risk systems must reach an appropriate level of accuracy, of resistance to errors and faults, and of cybersecurity, and must "perform consistently in those respects throughout their lifecycle" [1]. Article 15(3) says the levels of accuracy and "the relevant accuracy metrics" must be "declared in the accompanying instructions of use" [1].

The act sets no numeric accuracy level. Article 9(8) says the metrics and thresholds must be "appropriate to the intended purpose" of the system [1]. And "prior defined" means the threshold is written before the test, which is the practice described in using evals to decide a release.

When the EU AI Act's testing rules apply

As published in 2024, Article 113 of the act said the regulation "shall apply from 2 August 2026", with some chapters earlier and some later [1]. For the high-risk rules, that date has been replaced.

Regulation (EU) 2026/1744, called the Digital Omnibus on AI, is dated 8 July 2026 and was published on 24 July 2026 in the Official Journal, where EU laws are published [2]. It replaced point (c) of Article 113. The new text says that Chapter III, Sections 1, 2 and 3, with the exception of Article 6(5), apply from the two dates below. Articles 9 and 15 are in Section 2.

  • 2 December 2027 for AI systems classified as high-risk under Article 6(2) and Annex III, and
  • 2 August 2028 for AI systems classified as high-risk under Article 6(1) and Annex I [2].

The amending regulation gives its reason in recital 40, one of the numbered explanations at the start of the regulation: "the delayed availability of standards, common specifications, and alternative guidance and the delayed establishment of national competent authorities" [2]. In plain words, the supporting standards and the national authorities were late.

As of 30 September 2026, the European Commission's own AI Act Service Desk page for Article 113 carried a notice that its text had not yet been updated to reflect the amendments [3]. Which annex your system falls under, and so which date applies, is a legal classification for a lawyer to confirm.

What the act asks of companies that build the largest models

Chapter V of the act covers general-purpose AI models, the large models that other companies build products on. Article 55(1)(a) applies to providers of general-purpose models with systemic risk. It says they must "perform model evaluation in accordance with standardised protocols and tools reflecting the state of the art, including conducting and documenting adversarial testing of the model" [1]. Adversarial testing means testing by people or programs that try to make the model fail or misbehave. Chapter V has applied from 2 August 2025, and the 2026 amendment left that date unchanged [2].

The General-Purpose AI Code of Practice was published on 10 July 2025. The European Commission describes it as "a voluntary tool, prepared by independent experts", to help providers of these models comply with the act [4]. The Commission says its Safety and Security chapter is "only relevant to the small number of providers of the most advanced models" [4]. The chapters were not read for this page.

Article 55 is addressed to the provider of the model. A company that builds a product on another company's model should ask a lawyer which duties apply to it.

What NIST's two documents say

The National Institute of Standards and Technology (NIST) is a US government standards body. Its AI Risk Management Framework 1.0, released on 26 January 2023, describes itself as "voluntary" [5]. A customer contract can still ask a supplier to follow it.

The framework uses the abbreviation TEVV for test, evaluation, verification and validation, and says these tasks "are performed throughout the AI lifecycle" [5]. Its sentence on timing is: "AI systems should be tested before their deployment and regularly while in operation" [5]. Its MEASURE function lists outcomes that an organisation should be able to show [5]:

  • MEASURE 2.1: "Test sets, metrics, and details about the tools used during TEVV are documented."
  • MEASURE 2.3: performance is "demonstrated for conditions similar to deployment setting(s). Measures are documented."

The framework sets no pass mark, and it says of itself that "it does not prescribe risk tolerance" [5]. As of 30 September 2026, NIST's programme page says: "The AI RMF 1.0 is being revised as part of the White House AI Action Plan" [6]. The page gives no date for a new version.

NIST's Generative AI Profile, numbered AI 600-1 and released on 26 July 2024, is a companion document for systems that generate text, images or other content [7]. Two of its suggested actions are about testing [7].

  • MS-2.5-001: "Avoid extrapolating GAI system performance or capabilities from narrow, non-systematic, and anecdotal assessments." GAI is NIST's short form for generative AI. In plain words, a few good examples are too little to support a claim about the whole system.
  • GV-1.3-002: "Establish minimum thresholds for performance or assurance criteria and review as part of deployment approval".

The first of those actions is the reason a good AI demo does not show the product works.

What ISO/IEC 42001 is

ISO/IEC 42001:2023 is an international standard published jointly by ISO and the International Electrotechnical Commission (IEC), two international standards bodies. The IEC's catalogue lists it as edition 1.0, 51 pages, published on 18 December 2023. The catalogue says the document "specifies the requirements and provides guidance for establishing, implementing, maintaining and continually improving an AI (artificial intelligence) management system within the context of an organization" [8].

It is a standard for how an organisation manages AI. The full text is sold by the publishers and was not read for this page, so this page makes no statement about what the standard requires on testing. If a supplier cites it, ask to see the clause.

Stanford's 2026 AI Index reports that ISO/IEC 42001 was cited by 36% of respondents and the NIST AI Risk Management Framework by 33% [9]. The Index page that we read names no survey and gives no sample size, so use those figures only as evidence that both documents are in use.

What Colorado's law says

Colorado passed an AI law in 2024, and the version that will apply is a different one. The three facts below come from the Colorado General Assembly's own page for each bill, read on 30 September 2026.

  • SB24-205 was signed on 17 May 2024 with a start date of 1 February 2026. The legislature's summary listed duties that included an impact assessment and an annual review [12]. This version never applied.
  • SB25B-004, signed on 28 August 2025, moved the start date to 30 June 2026 [13].
  • SB26-189, signed on 14 May 2026, repealed and reenacted those provisions "with new requirements regarding the use of automated decision-making technology in consequential decisions". Its duties start on 1 January 2027 [11].

According to the General Assembly's summary of SB26-189, a developer must give a business that uses its technology "technical documentation describing the covered ADMT's intended uses, categories of training data, known limitations, and instructions for appropriate use and human review". ADMT is the act's short form for automated decision-making technology. Developers and deployers (the act's word for businesses that use the technology) must keep the records needed to demonstrate compliance "for at least 3 years", and consumers can request human review after an adverse decision.

The summary uses none of the words test, evaluation or impact assessment. So the Colorado law that starts in 2027, as the legislature summarises it, states no testing duty. The bill text itself was not read for this page. A developer does have to describe "known limitations", and a test is the ordinary way to find one. That last point is our reading, and a lawyer should confirm it.

What one US enforcement action alleged about testing

The US Federal Trade Commission (FTC) enforces consumer protection law, and it has applied that law to an AI product. In September 2024 it announced a complaint against DoNotPay, which alleged that the company "did not conduct testing to determine whether its AI chatbot's output was equal to the level of a human lawyer" [10]. The company settled, so no court ruled on the allegation. That case and others are in public AI failures and the tests that target each one.

Whether eval results count as compliance evidence

The texts on this page describe records, and an eval run produces a record of that type. NIST's MEASURE 2.1 asks that test sets and metrics "are documented" [5]. EU AI Act Article 9(8) describes testing against thresholds defined in advance, and Article 15(3) asks for accuracy levels to be declared [1].

An eval run that can serve as such a record has five parts: the test set and where its cases came from, the threshold and the date it was written, the version of the model and the product that was tested, the result with the failures listed, and the role that approved the release. How to read an AI eval report covers what a report should contain.

Whether a given record satisfies a given legal duty is for a lawyer to decide. For regulated sectors, see evals as compliance evidence in our financial software guide, the regulated industry software guide, and our post on how to negotiate a software contract you can verify.

How Reveneau applies this

Reveneau is an AI software development consultancy. All of our code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before it is released. The texts on this page describe the same order: thresholds defined before the test, and testing before release.

For a product in a regulated area, we ask every client at the start which laws and standards their lawyer has said apply, and we write the expectations and thresholds in the eval suite based on those texts. Legal advice, and the decision on which law applies, stay with the client's lawyer. Our part is the tests and the dated record of each run.

The checks that need judgment are graded by Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader. Reveneau, as a company, takes responsibility for the whole project through production and after release.

To discuss such a product, see AI development at Reveneau or contact us. The reasons for testing AI products in general are in why AI evals matter.

Common questions

Does the EU AI Act require testing?

Yes, for the systems the EU AI Act classifies as high-risk. Article 9(6) says high-risk AI systems shall be tested, and Article 9(8) says the testing happens before the system is placed on the market, against metrics and thresholds defined in advance. Most business software is outside the high-risk category, and a lawyer should confirm how your product is classified.

When does the EU AI Act apply to high-risk AI systems?

As of 30 September 2026, the EU AI Act's high-risk rules apply from 2 December 2027 for systems classified under Annex III and from 2 August 2028 for systems classified under Annex I. Regulation (EU) 2026/1744, published on 24 July 2026, set those dates. The date in the act as first published was 2 August 2026, so older documents can be wrong.

What is the NIST AI RMF?

The NIST AI RMF is the AI Risk Management Framework, a document from the US National Institute of Standards and Technology released on 26 January 2023. It is a voluntary guide to managing the risks of AI systems. On testing, it says AI systems should be tested before their deployment and regularly while in operation, and that test sets and metrics should be documented.

Is the NIST AI RMF mandatory?

The NIST AI RMF is voluntary. The framework describes itself as voluntary. It sets no pass mark, and it says of itself that it does not prescribe risk tolerance. A customer contract can still ask a supplier to follow it. As of 30 September 2026, NIST's programme page says version 1.0 is being revised and gives no date for the new version.

What is ISO 42001?

ISO/IEC 42001:2023 is an international standard for an AI management system, which means how an organisation manages its use of AI. The IEC catalogue lists it as edition 1.0, 51 pages, published on 18 December 2023. The text is sold by the publishers and was not read for this page, so ask any supplier who cites it to show the clause on testing.

Do evals count as compliance evidence?

Eval results are the type of record that these texts describe, and a lawyer decides whether they meet a specific duty. NIST's MEASURE 2.1 asks that test sets and metrics are documented, and EU AI Act Article 9(8) describes testing against thresholds defined in advance. Keep the test set, the dated threshold, the tested version, the result and the approving role together.

Does Colorado's AI law require testing?

Colorado's current AI law, SB26-189, states no testing duty in the General Assembly's summary. The summary asks developers for technical documentation, including known limitations, asks developers and the businesses that use their technology to keep records for at least 3 years, and gives consumers a right to request human review. Its duties start on 1 January 2027. The bill text was not read for this page.

Does any regulator say what accuracy is enough for an AI system?

None of the texts on this page gives a number for enough accuracy. The EU AI Act asks for an appropriate level of accuracy and for the accuracy metrics to be declared in the instructions of use, under Article 15. NIST's framework says it does not prescribe risk tolerance. The organisation sets its own threshold and writes it down before the test.

Which EU AI Act testing rules cover a company that builds on another company's model?

Article 55 of the EU AI Act, on model evaluation and adversarial testing, is addressed to providers of general-purpose AI models with systemic risk, meaning the companies that build the largest models. Articles 9 and 15 cover high-risk AI systems. A company that builds a product on another company's model should ask a lawyer which duties apply to it.

What should I do if a document still gives 2 August 2026 for the high-risk rules?

Treat a document that gives 2 August 2026 for the EU AI Act's high-risk rules as out of date and check when it was written. Regulation (EU) 2026/1744 replaced that date on 24 July 2026. As of 30 September 2026, even the European Commission's AI Act Service Desk page for Article 113 carried a notice that its text had not been updated.

Why did the EU move the dates for the high-risk rules?

The EU moved the dates because the supporting standards and the national authorities were late. Recital 40 of Regulation (EU) 2026/1744 gives the reason as the delayed availability of standards, common specifications and alternative guidance, and the delayed establishment of national competent authorities. The rules for general-purpose AI models in Chapter V kept their date of 2 August 2025.

Has a US regulator acted against a company for not testing its AI?

Yes, in one case described on this page, as an allegation that the company settled. The US Federal Trade Commission enforces consumer protection law. Its complaint against DoNotPay, announced in September 2024, alleged that the company did not conduct testing to determine whether its chatbot's output was equal to the level of a human lawyer. The company settled, so no court ruled on the allegation.

References