What goes wrong without evals

What happens to your product when the AI model changes

Every change a vendor makes to an AI model reaches the products built on that model. As of 30 September 2026, OpenAI states at least 6 months of notice before retiring a generally available model, Anthropic states at least 60 days, and Google states no fixed period. A 2023 paper found GPT-4's accuracy on one task fell from 84.0% to 51.1% between two versions three months apart, while another task improved. Pin the model version, record its retirement date, and keep an eval set you can run again. With that set, a model change is a re-run and a reading of the failures. Without it, the first report of a problem comes from a customer.

Published September 30, 2026. Editorial.

Key takeaways

  • As of 30 September 2026, OpenAI states at least 6 months of notice before retiring a generally available model and Anthropic states at least 60 days. Google's Gemini page states no fixed period.
  • Chen, Zaharia and Zou measured GPT-4 on 1,000 prime number questions: 84.0% correct in March 2023 and 51.1% in June 2023. On 7,405 multi-step questions the same model rose from 1.2% to 37.8%.
  • Pin a dated model version so behaviour changes only on a day you choose, and record the retirement date, because dated versions are the ones vendors retire.
  • A regression is a task the product handled before a change and fails after it. Re-running a fixed eval set on the old and new model lists every regression case by case.
  • Keep the eval set in files your company controls: OpenAI's own online Evals platform becomes read-only on 31 October 2026 and shuts down on 30 November 2026.

On 30 September 2026 Anthropic notified developers that Claude Sonnet 4.5 will be retired from its API on 30 November 2026 [2]. An API (application programming interface) is the connection a product uses to send text to a model and get text back. From the notice to the retirement is 61 days: 31 in October and 30 in November. After that date, in Anthropic's words, "Requests to retired models will fail" [2].

Any product that calls that model through Anthropic's API has 61 days to move to another model and to show that it still works. This page explains how often that happens, what can change even when no model is retired, and how a set of evals makes a model change a planned task with a date. An eval is a repeatable test: a set of inputs, a written expectation for each one, and a way to score the output.

Vendors retire models on their own schedule

Each large model vendor publishes a page of retirement dates. The table shows what each page stated on 30 September 2026. Every line is a vendor's statement of its own policy, and a vendor can change it.

Vendor Stated notice Example on the vendor's page Days between the dates
OpenAI [1] "At least 6 months" for generally available models, "At least 3 months" for specialised variants Notice 11 June 2026, removal 11 December 2026 183
Anthropic [2] "at least 60 days' notice" for publicly released models Notice 30 September 2026, retirement 30 November 2026 61
Google, Gemini API [3] No fixed period is stated gemini-2.0-flash released 5 February 2025, shutdown date 1 June 2026 481

Generally available means released for all customers to use. The day counts are this page's arithmetic on the dates each vendor prints. The Google row measures from release to shutdown, because Google's page gives no notice date.

Three details matter for planning. OpenAI says preview models, which have "preview" in their name, "may be retired with much shorter notice, such as 2 weeks". It advises against using them for work a business depends on, unless the customer can move to another model at short notice [1]. Google says its listed dates are "the earliest possible dates on which a model might be retired" and promises "advance notice" without a number [3]. Anthropic says that Amazon Bedrock and Google Cloud "set their own retirement schedules" [2], so the date depends on where your product reaches the model. On what retirement means, OpenAI's page says the model "will no longer be accessible" [1], and Google's says a model that is shut down "is completely turned off" [3].

How long does a model version stay available? Three examples from these pages give 365, 481 and 491 days. Claude Opus 4.1 has the date 5 August 2025 in its name and was retired on 5 August 2026 [2]. Gemini 2.0 Flash is the 481-day row above [3]. The OpenAI version named gpt-5-2025-08-07 is scheduled for removal on 11 December 2026 [1], which is 491 days after the date in its name. Other models last longer: Google lists gemini-2.5-pro, released 17 June 2025, with "No shutdown date announced" [3].

A model can change behaviour between versions

Retirement is the visible change. The less visible one is that a newer version of the same model answers differently. The best documented case is a 2023 paper by Lingjiao Chen, Matei Zaharia and James Zou, published as a preprint, which is a research paper released before formal review by a journal [4]. They ran the same questions on the March 2023 and June 2023 versions of GPT-4, a large language model (LLM) from OpenAI.

Task Questions GPT-4, March 2023 GPT-4, June 2023
Say whether a number is prime 1,000 84.0% correct 51.1% correct
Write code that runs as written 50 52.0% 10.0%
Answer a sensitive question 100 21.0% answered 5.0% answered
Multi-step questions, exact match 7,405 1.2% 37.8%
Medical exam questions 340 86.6% 82.4%

The table shows changes in both directions. Accuracy on prime numbers fell by 32.9 percentage points (84.0 minus 51.1), and the exact match rate on multi-step questions rose by 36.6 points (37.8 minus 1.2). The older GPT-3.5 model changed in the opposite direction on prime numbers, from 49.6% to 76.2% [4]. The right word is "changed". A new version can be better on some tasks and worse on others, and only a test on your own task shows which.

Two limits. These are models from 2023, and nothing in the paper describes a current model. Part of the fall in the code row came from the model adding formatting marks around its code, which the authors' check counted as code that does not run as written [4]. That detail is itself a lesson: a change in the format of the output can stop the software that reads it from working.

The authors conclude that their results show "the need for continuous monitoring of LLMs" [4].

Every change to the model reaches your product

A product built on a model has three parts: your prompt (the written instructions your product sends to the model), your documents, and the vendor's model. You control the first two. When the third changes, every answer your product gives can change with it, including answers to questions that worked for a year.

Length can change as well as correctness. In the prime number task, the average GPT-4 answer went from 638.3 characters in March 2023 to 3.9 characters in June 2023 [4]. A screen designed for a paragraph now shows four characters.

Pinning a version delays the change and sets a deadline

Model names on the vendor pages include dates, such as gpt-5-2025-08-07 and claude-opus-4-1-20250805 [1] [2]. OpenAI calls these dated versions snapshots. Telling your product to use one dated version is called pinning.

Pin the version. The aim is that your product changes behaviour on a day you choose, after a test. Pinning has a limit: the dated versions are the ones that get retired. The pinned version gpt-5-2025-08-07 has a removal date of 11 December 2026 [1]. So pin, and put the retirement date in the calendar on the day the notice arrives.

A regression is something that worked and then stopped working

A regression is a task the product handled before a change and fails after it. Anthropic's engineering blog describes regression evals as the ones that ask "Does the agent still handle all the tasks it used to?" and says their pass rate should stay at or just under 100 percent [5].

Without an eval set, the only way to find a regression is a customer complaint. The same post says: "Absent evals, debugging is reactive: wait for complaints, reproduce manually, fix the bug, and hope nothing else regressed" [5]. Debugging means finding and fixing faults.

With an eval set, the question "does it still work?" is answered by running the set again. If the set runs in minutes, the run and a reading of the failures fit in one afternoon. OpenAI's documentation makes the same point about timing: writing evals matters "especially when upgrading or trying new models" [6]. If you have no set yet, what an AI eval is is the place to start.

What to re-run and when

What changed What to re-run When
A new model version you choose to adopt The whole eval set, on the old and the new version Before the switch
A retirement notice from the vendor The whole set, on the replacement the vendor names The week the notice arrives
A new prompt The whole set Before the prompt is released
A new document source The cases that depend on documents, plus the source checks Before the source is in use
Nothing you know of The whole set On a fixed date each month

The last row exists because of the 2023 paper, which describes the behaviour of the same service changing between March and June [4]. NIST, the United States standards body, names the "variability of GAI system performance over time" in its suggested actions, where GAI is its short form for generative AI [7].

How to test a model upgrade

  1. Keep the eval set fixed. Add and edit nothing during the comparison.
  2. Run the set on the current model. Record the pass rate and the list of failed cases. Run each case more than once, because output varies between runs.
  3. Run the same set on the new model with the same prompt.
  4. Compare case by case. List every case that passed before and fails now. These are the regressions.
  5. Read every regression. One failure on a case that blocks release matters more than the average.
  6. Check answer length, format, response time and cost per answer.
  7. Decide against the pass line you set before the run, as using evals to decide a release describes.
  8. Keep both result files, with the model name and the date.

A fall of two points between two runs can be inside the range that a small set produces by chance. How to read an AI eval report shows the arithmetic. If you are comparing several candidate models, use choosing a model with your own evals.

Keep the eval set in files you own

Testing tools get retired too. OpenAI's documentation says: "OpenAI is deprecating the Evals platform." Deprecating means announcing that a product will be withdrawn. Existing evals become read-only on 31 October 2026 and the platform is scheduled to shut down on 30 November 2026 [6]. OpenAI's model retirement page points to a guide for moving to another tool [1]. The shutdown concerns OpenAI's own online tool. The practice of writing evals continues.

Store the inputs, the expectations and the past results in files your company controls, in a format any tool can read. Ownership of the set is covered in what AI evals cost and who owns them.

What to ask a vendor or a development partner

  • Which exact model version does the product call today?
  • Who reads the model vendor's retirement page, and how often?
  • How many working days pass between a notice and a tested switch?
  • Where is the eval set stored, and who owns it when the contract ends?

How to evaluate a vendor's eval suite gives more questions of this type, and why AI evals matter gives the wider argument.

How Reveneau handles a model change

At Reveneau all code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before it is released. We treat a new model version as a change like any other: it must pass the same suite before it replaces the version that customers use. We pin the model version in each product, record the vendor's retirement date, and plan the re-run as dated work inside the notice period.

The judged checks in our suite are graded with Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader. The grader is a model with versions of its own, so we re-check it against human labels when its version changes, as calibrating Jev against human labels describes.

Reveneau, as a company, takes responsibility for the whole project through production and after release, and a model retirement after release is part of that. See AI development at Reveneau or contact us.

Best for

  • Product owners whose product calls a model from OpenAI, Anthropic, Google or a cloud platform
  • Teams that have received a retirement notice and need a plan with dates
  • Buyers who want to ask a development partner how model changes are handled

Avoid if

  • You are choosing a model for a new product: use the guide on choosing a model with your own evals
  • You need the method for building the eval set itself, which the LLM evals guide covers

Check before you decide

  • Re-read each vendor's retirement page before you act: the dates here were read on 30 September 2026
  • Confirm which exact model version your product calls today
  • Check whether a cloud platform between you and the vendor sets its own retirement date

Common questions

How often do AI models get retired?

AI model versions stayed available for 365 to 491 days in the three examples this page takes from the vendors' own pages, read on 30 September 2026. Claude Opus 4.1 was retired 365 days after the date in its name. Gemini 2.0 Flash has a shutdown date 481 days after release. The OpenAI version gpt-5-2025-08-07 is scheduled for removal 491 days after the date in its name. Other models last longer.

How much notice does a vendor give before retiring an AI model?

Notice before an AI model is retired depends on the vendor and the type of model. As of 30 September 2026, OpenAI's documentation states at least 6 months for generally available models, at least 3 months for specialised variants, and as little as 2 weeks for preview models. Anthropic's documentation states at least 60 days. Google's Gemini page states no fixed number and lists the earliest possible shutdown dates.

Can a model update change how my product behaves?

Yes. A model update can change both the correctness and the form of your product's answers. Chen, Zaharia and Zou compared the March 2023 and June 2023 versions of GPT-4 on 1,000 prime number questions and measured 84.0% correct, then 51.1%. The average answer length on that task went from 638.3 characters to 3.9. On 7,405 multi-step questions the same model improved from 1.2% to 37.8%.

Should I pin a model version in my product?

Yes, pin a dated model version, and treat its retirement date as a deadline. A pinned version keeps your product's behaviour the same until a day you choose, after a test. Pinning does not last for ever, because dated versions are the ones vendors retire: OpenAI's page lists gpt-5-2025-08-07 for removal on 11 December 2026. Put that date in the calendar when the notice arrives.

What is a regression in an AI product?

A regression in an AI product is a task the product handled before a change and fails after it. The change can be a new model version, a new prompt or a new document source. Anthropic's engineering blog describes regression evals as the tests that ask whether the product still handles all the tasks it used to, and says their pass rate should stay at or just under 100 percent.

How do I test a model upgrade before I switch?

To test a model upgrade, keep your eval set fixed, run it on the current model and record the result, then run the same set on the new model with the same prompt. Compare case by case and list every case that passed before and fails now. Read each of those failures, check answer length, format, speed and cost, and decide against the pass line you set before the run.

What happens to my product if I do nothing when a model is retired?

If you do nothing, the product stops working on the retirement date. Anthropic's documentation says requests to retired models will fail, OpenAI's says the model will no longer be accessible, and Google's says the model is completely turned off. A product that still calls the retired name on that day returns errors to its users until somebody changes the model name and releases the change.

Is a newer AI model always more accurate on my task than the one it replaces?

No. A newer model can be more accurate on some tasks and less accurate on others. In the 2023 paper by Chen, Zaharia and Zou, the June version of GPT-4 was 32.9 percentage points less accurate than the March version on 1,000 prime number questions and 36.6 points more accurate on 7,405 multi-step questions. Only a test on your own task shows how your product changes.

Do I need to re-run my evals when only the prompt changes?

Yes. Re-run the whole eval set when the prompt changes, because the prompt is the set of written instructions sent with every request, so a change to it can affect every case. Run the set before the new prompt is released, compare the result case by case with the last run, and read each case that passed before and fails now. The same rule applies to a new document source.

How long does it take to re-test a product on a new model?

Re-testing a product on a new model takes as long as the eval set takes to run, plus the time to read the failures. If the set runs in minutes, both fit in one afternoon. No source on this page measures an average. The time that matters is the notice period, which runs from 60 days at Anthropic to 6 months at OpenAI for generally available models, as of 30 September 2026.

Do retirement dates differ if my product reaches the model through a cloud platform?

Yes, retirement dates can differ on a cloud platform. Anthropic's documentation says that partner-operated platforms, naming Amazon Bedrock and Google Cloud, set their own retirement schedules, so a model's status and dates can differ from those on Anthropic's own service. Find out which platform your product calls, read that platform's retirement page as well as the model vendor's, and record the earlier of the two dates.

What happens to evals stored in OpenAI's Evals platform?

Evals stored in OpenAI's own online Evals platform become read-only on 31 October 2026, and the platform is scheduled to shut down on 30 November 2026, according to OpenAI's documentation read on 30 September 2026. OpenAI's model retirement page points to a guide for moving to another tool. The practice of writing evals continues. Store your inputs, expectations and past results in files your company controls.

References