Why AI evals matter: the test that shows whether an AI product works
An AI product can give different output on different days and for different inputs, so a demo, a vendor's promise and a public benchmark score are weak evidence that it works for your customers. An AI eval is the stronger evidence: a written, repeatable test with an input, an expectation and a scoring rule. This guide is written for the people who approve budgets and releases. It explains what an eval is, what public records show about wrong AI answers and advertising claims, how to read a pass rate and use it to decide a release, what the work costs and who owns it, and what the texts of regulators say about testing.
Published September 30, 2026. Editorial.
Key takeaways
- An AI eval is a repeatable test: an input, a written expectation of a correct output, and a rule that scores the output. The full set for one product is an eval suite.
- A demo is one run on a chosen input. A 2024 study of five language models found accuracy differences of up to 15% between runs, under settings meant to make output repeatable.
- On February 14, 2024, a British Columbia tribunal ordered Air Canada to pay $812.02 after finding it did not take reasonable care to ensure its chatbot was accurate.
- The AI model your product uses can change. On 1,000 prime-number questions, GPT-4's accuracy was 84.0% in March 2023 and 51.1% in June 2023, as measured by Chen, Zaharia and Zou.
- A 92% pass rate on 50 cases has a 95% range of 84.5% to 99.5%. On 1,000 cases the range is 90.3% to 93.7%. Always ask for the number of cases.
- Write the pass line before the test runs. NIST's generative AI profile suggests minimum thresholds reviewed as part of deployment approval, and NIST leaves the number to each organisation.
- As of 30 September 2026, the EU AI Act's testing rules for high-risk systems apply from 2 December 2027 or 2 August 2028, depending on the type of system. A lawyer should confirm what applies to you.
On February 14, 2024, the British Columbia Civil Resolution Tribunal ordered Air Canada to pay a customer, Jake Moffatt, a total of $812.02 in Canadian dollars. A chatbot on the airline's website had told Moffatt they could apply after travelling for a bereavement fare, a reduced fare for people who travel because of a death in the family. A page on the same website said the opposite. The tribunal member wrote: "I find Air Canada did not take reasonable care to ensure its chatbot was accurate" [1].
The amount is small, and a small claims decision binds only the two parties. The reasoning is what a business owner should note. The tribunal treated the chatbot as part of the company's website and wrote: "It makes no difference whether the information comes from a static page or a chatbot" [1].
This guide is written for the person who approves the budget or the release of an AI product and who has never written a test. An AI product gives different output on different days and on different inputs. A demo, a vendor's promise and a public benchmark score (a model's result on a shared public test) are therefore weak evidence that the product works for your customers, and a written, repeatable test, called an eval, is the evidence to ask for.
An eval is a written test that can be run again
An AI eval is a repeatable test of an AI product. It has three parts: an input (the question or document the product receives), a written expectation of what a correct output must do, and a rule that scores the output, most often as pass or fail. The full set of these cases for one product is called an eval suite, and the share of cases that pass is the pass rate.
OpenAI's documentation gives the reason ordinary software tests are insufficient here: "Generative AI is variable. Models sometimes produce different output from the same input, which makes traditional software testing methods insufficient for AI architectures" [2]. An ordinary program gives the same output for the same input every time, so one check is enough. A language model, the type of AI model that produces text, can answer the same question correctly on Monday and wrongly on Tuesday. Each case is therefore run several times, across many cases, and the result is a rate.
The expectations come from people who know the business. Take an invented furniture shop whose policy allows returns for 30 days. The shop knows that rule and the AI model has to be told it, so the shop is the one who can say that an answer promising 90 days is a failure. Engineers then turn each expectation into a check a computer can run. The page what is an AI eval shows one case written out in a table and explains the words a vendor will use.
A good demo is one run on a chosen input
A demo shows that the product can give a good answer once, on an input the presenter selected and rehearsed. It gives no information about how often the product is wrong, because each input ran one time.
Research posted in August 2024 by Berk Atil and 12 other authors shows how much one run can mislead. They took five large language models, configured them with the settings meant to give the same output every time, and ran eight common tasks 10 times each. They report "accuracy variations up to 15% across naturally occurring runs" [3]. They also report that "none of the LLMs consistently delivers repeatable accuracy across all tasks" [3]. LLM is short for large language model. That is the largest difference they observed, in a paper posted before formal review. If accuracy can move that much between runs under settings chosen to prevent it, a single run in a meeting is a weak basis for a purchase.
NIST, the United States standards institute, gives the same advice in its 2024 profile for generative AI: "Avoid extrapolating GAI system performance or capabilities from narrow, non-systematic, and anecdotal assessments" [4]. GAI is NIST's short form for generative AI.
A buyer may also ask what share of AI pilots reach real use by customers. This guide gives no percentage, because no measured figure could be verified. After a demo, ask in writing for the eval set, the pass rate with the number of cases it was measured on, and the failed cases. The page why a good AI demo does not show the product works has the full argument and a table comparing what a demo and an eval each show.
Public records show what a wrong answer or claim has cost
Tribunals and regulators have already acted on wrong answers from AI products and on advertising claims about them. In the Air Canada decision above, the order of $812.02 was made up of $650.88 in damages, $36.14 in interest and $125 in tribunal fees [1].
In the United States, the Federal Trade Commission (FTC) finalized an order against DoNotPay on February 11, 2025. The FTC's release says the company had promoted its service as "the world's first robot lawyer". The order requires it to pay $193,000 and prohibits it from advertising that the service "performs like a real lawyer" unless it has sufficient evidence for the claim [5]. It is a consent order, which means the company settled and no court ruled on the facts.
A second FTC case shows the difference between a claimed figure and a tested one. According to the FTC's complaint, Workado promoted its AI content detector as "98 percent" accurate, and independent testing showed an accuracy rate of 53 percent on general-purpose content [6]. That case was also settled, and the FTC's release leaves out who ran the test and how many texts it used.
Each of these records concerns the accuracy of a statement: an answer the product gave, or a claim the company made about the product. A test written in advance, with its results kept, is how a company produces evidence of accuracy. The page public AI failures and the tests that target each one describes each record and names the type of test case that targets each type of failure.
The AI model your product uses changes on the vendor's schedule
Most AI products are built on a model that another company owns, updates and retires. Each of those changes reaches the product.
Lingjiao Chen, Matei Zaharia and James Zou measured one such change. They gave the March 2023 and June 2023 versions of GPT-4 the same 1,000 questions asking whether a number is prime. They report: "GPT-4's accuracy dropped from 84.0% in March to 51.1% in June", while GPT-3.5 on the same questions rose from 49.6% to 76.2% [7]. The changes went in both directions, and the figures describe 2023 models only. What still applies is the authors' conclusion: a service that keeps the same name can behave differently a few months later, so it needs continuous monitoring.
Vendors also retire models. As of 30 September 2026, Anthropic's documentation says the company gives "at least 60 days' notice before model retirement for publicly released models", and that "Requests to retired models will fail" [8]. That is a company's statement of its own policy, which it can change.
Without an eval suite, the first sign of a change for the worse is a customer complaint. With one, a model change means running the same cases again and comparing the two results before customers see the new model. The page what happens to your product when the AI model changes lists what to run again and when.
A pass rate means little without the number of cases it was measured on
An eval report gives a pass rate. The first question to ask is how many cases it was measured on, because a small set leaves a wide range of possible true rates.
Anthropic's explanation of its statistics paper gives the method. A 95% confidence interval, the range that is likely to contain the true rate, "can be calculated from the SEM by adding and subtracting 1.96 × SEM from the mean score" [9]. SEM is the standard error of the mean, a number that says how far the score could move if a different set of questions had been drawn. For a pass rate p measured on n cases, the standard formula is the square root of p x (1 - p) / n. Here is the calculation for a pass rate of 92%, as an illustration.
| Cases | Standard error | 1.96 x standard error | 95% range |
|---|---|---|---|
| 50 | square root of (0.92 x 0.08 / 50) = 0.0384 | 0.0752 | 84.5% to 99.5% |
| 1,000 | square root of (0.92 x 0.08 / 1,000) = 0.0086 | 0.0168 | 90.3% to 93.7% |
On 50 cases, a reported 92% could be a true rate as low as 84.5%. On 1,000 cases, the lower end of the range is 90.3%.
An average can also hide a group that fails. In the 2018 Gender Shades study, Joy Buolamwini and Timnit Gebru tested 3 commercial systems that classify gender from a photograph of a face. The error rate was up to 34.7% for darker-skinned women, and the maximum for lighter-skinned men was 0.8% [10]. The study is about face analysis software, and it shows one overall score hiding a group that fails. Ask for results split by the groups of customers and types of request that matter to you. The page how to read an AI eval report is a printable checklist for this.
Set the pass line before the test is run
An eval result is useful when it decides something. The pass line is the lowest pass rate at which you will release. NIST's generative AI profile suggests this action: "Establish minimum thresholds for performance or assurance criteria and review as part of deployment approval" [4]. A threshold is a minimum value, and deployment is NIST's word for release. NIST's AI Risk Management Framework describes the decision itself: "A determination is made as to whether the AI system achieves its intended purposes and stated objectives and whether its development or deployment should proceed" [11]. NIST leaves the number to you: the framework says it "does not prescribe risk tolerance" [11]. Both documents are voluntary.
Write the pass line down before the test runs, because a line chosen after the result is known can be set to match the result. Set a higher line for outputs that can cost a customer money or cause harm, and name the failures that block a release whatever the average is. For the invented furniture shop, one such failure is any answer that promises a refund outside the policy. And make the release a decision of the company, recorded together with the result it was based on. A check that a change must pass before release is called a release gate.
When a failure is found after release, add it to the suite as a permanent case, so every later change is tested for it. The page using evals to decide whether an AI feature is ready to release gives a table by type of output.
What evals cost and who owns them
No source we could verify measures what companies spend on evals, in money or in staff time, so this guide gives no typical budget. The cost has four parts: writing the expectations, collecting and labelling cases, running the tests, and keeping the set current.
One published figure shows the running cost of a large research project. In October 2025, Sayash Kapoor, Benedikt Stroebl and 29 other authors reported 21,730 agent attempts across 9 models and 9 public benchmarks, at a total cost that the authors round to $40,000 [12]. An agent is an AI program that takes steps on its own. That cost belongs to one study of public models, and a company's suite for one product is a different size.
On where to start, Anthropic's engineering post says "20-50 simple tasks drawn from real failures is a great start" and warns that "Evals get harder to build the longer you wait" [13]. Both statements are the company's advice from its own work.
Ownership has three parts. The people responsible for the product own what a correct answer is. Engineering owns running the tests. The company owns the release decision. NIST's framework adds, in its MEASURE 1.3 outcome, that regular assessments should involve internal experts who did not build the system, or independent assessors [11]. If a development partner builds the product, agree in the contract that the eval set, with its cases and expectations, belongs to you when the contract ends. The page what AI evals cost and who should own them sets out each cost part and who does it.
What regulators and standards bodies say about testing
In the European Union's AI Act, Article 9(8) says of high-risk AI systems: "Testing shall be carried out against prior defined metrics and probabilistic thresholds that are appropriate to the intended purpose of the high-risk AI system" [14]. That means deciding what you will measure, and the minimum result, before you test. The article applies to systems the Act classes as high-risk, a legal category that most business software is outside.
The dates changed in 2026. As of 30 September 2026, Regulation (EU) 2026/1744, published on 24 July 2026, makes the high-risk rules apply from 2 December 2027 for systems classed as high-risk under Annex III of the Act, and from 2 August 2028 for those classed under Annex I. The date it replaced was 2 August 2026 [15].
In the United States, the NIST AI Risk Management Framework is voluntary. It says: "AI systems should be tested before their deployment and regularly while in operation" [11].
This section describes what the texts say and is general information. A lawyer should confirm what applies to a specific product. The page what regulators and standards bodies say about testing AI covers each text with its article numbers and dates, and the EU AI Act and your product explains the Act's risk classes.
What to read after this guide
If you read one more page, read the one on eval reports and take its checklist to your next vendor meeting. For the questions to ask about the suite itself, see how to evaluate a vendor's eval suite. Teams that will build the tests can continue with the method guides: LLM evals, AI agent evals and AI benchmarks and your own evals. Evals for code that AI writes are covered in eval-driven development.
Where Reveneau fits
Reveneau is an AI software development consultancy, and this guide describes how we work. All of Reveneau's code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before release. The specification is the written description of what the software must do. For an AI product, we write the expected behaviour for each case with the client, set the pass line before the test runs, and show the client the cases, the pass rate and the failures.
The checks in that suite that need judgement are graded by Jev, TypeSafe AI's decision model, which returns a probability for a written question instead of writing text. On Reveneau's own suite the run is ten times faster than with our previous language-model grader, so the suite can run on every change.
We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. Reveneau, as a company, takes responsibility for the whole project through production and after release. To plan an AI product with its tests written first, see AI development at Reveneau or contact us.
Explore the guide
Start here
What is an AI eval? A plain-language explanation
An AI eval is a repeatable test of an AI product. It has three parts: an input, a written expectation of what a correct output must do, and a rule that scores the output as pass or fail. A language model can give a different answer to the same question on a different day, so each case is run several times, and the full set, called an eval suite, reports a pass rate. OpenAI's documentation says this variation makes traditional software testing methods insufficient for AI. People who know the business write the expectations, and engineers turn them into checks a computer can run.
Why a good AI demo does not show the product works
A demo is one run of an AI product on an input the presenter chose, so it shows that the product can give a good answer once. Three things limit what it proves: the inputs were selected, a language model's output changes between runs, and real customers send inputs nobody planned for. In a 2024 study of five language models set up to give the same output every time, accuracy still varied by up to 15% between runs across eight tasks. After a demo, ask for three things in writing: the eval set, the pass rate with the number of cases it was measured on, and the failed cases.
What goes wrong without evals
Public AI failures and the tests that target each one
Six public records show four types of AI failure that a test can target. A tribunal ordered Air Canada to pay $812.02 after its chatbot stated a rule the airline's policy page contradicted. A court imposed $5,000 on lawyers who filed invented court opinions. The Markup reported a New York City chatbot giving answers that conflict with city law. Regulators challenged advertising claims by Workado, DoNotPay and Pieces. A written test, called an eval case, targets each type: policy questions scored against the policy text, a lookup of every cited source, legal questions with expectations written by a specialist, and a test set built from the content real customers send.
What happens to your product when the AI model changes
Every change a vendor makes to an AI model reaches the products built on that model. As of 30 September 2026, OpenAI states at least 6 months of notice before retiring a generally available model, Anthropic states at least 60 days, and Google states no fixed period. A 2023 paper found GPT-4's accuracy on one task fell from 84.0% to 51.1% between two versions three months apart, while another task improved. Pin the model version, record its retirement date, and keep an eval set you can run again. With that set, a model change is a re-run and a reading of the failures. Without it, the first report of a problem comes from a customer.
Decide with evals
How to read an AI eval report if you are not an engineer
A pass rate means little until you know what it was measured on. A result of 92 percent on 50 cases has a 95 percent confidence interval of 84.5% to 99.5%. The same 92 percent on 1,000 cases has an interval of 90.3% to 93.7%. To read an AI eval report, check six things: the number of cases and where they came from, the range around the pass rate, the pass rate for each group of cases, the list of failures, who or what did the grading and how many times each case ran, and whether the same set was run on the previous version. This page shows the arithmetic and ends with a checklist to print.
Using evals to decide whether an AI feature is ready to release
An AI feature is ready to release when it meets a pass line that the company wrote down before the test was run. Set one line for the whole test set, a higher line for outputs that can cost money or cause harm, and a short list of failures that block release whatever the average is. The company makes the decision and records it, and the people who run the tests report the result without owning the decision. Every failure found after release becomes a permanent test case. The EU AI Act describes testing against thresholds defined in advance, NIST's Generative AI Profile describes minimum thresholds reviewed when deployment is approved, and both leave the number to you.
Budget and rules
What AI evals cost and who should own them
No public source measures what a typical company spends on AI evals, so this page gives you the parts to price yourself. The cost has four parts: writing down what a correct output is, collecting and labelling test cases, paying for the model usage each test run consumes, and keeping the set current as the product and the model change. Three of those are people's time. The running cost is arithmetic on a vendor's published price. Ownership follows the same split: the product side owns what correct means, engineering owns the running, and the company owns the release decision. When a partner builds the product, the contract should say that you keep the test set.
What regulators and standards bodies say about testing AI
As of 30 September 2026, the EU AI Act states a testing duty for AI directly: Article 9 requires high-risk AI systems to be tested before they are placed on the market, against metrics and thresholds defined in advance. After an amendment published on 24 July 2026, that duty applies from 2 December 2027 or 2 August 2028, depending on the type of system. In the United States, NIST's AI Risk Management Framework describes testing before release and during operation, and it is voluntary. ISO/IEC 42001 is a management standard, and Colorado's 2026 law asks for documentation and records. This page describes what the texts say. It is general information, and a lawyer should confirm what applies to your product.
Common questions
Why do AI evals matter to a business?
AI evals matter to a business because an AI product can answer the same question correctly one day and wrongly the next, and customers act on the answers. In February 2024 a British Columbia tribunal ordered Air Canada to pay a customer $812.02 after finding the airline did not take reasonable care to ensure its chatbot was accurate. An eval is a written, repeatable test of what the product tells customers.
Do I need evals if I buy an AI product instead of building one?
Yes. A buyer needs the vendor's eval results for the same reason a builder needs evals: a demo is one run on a chosen input. Ask the vendor in writing for the eval set, the pass rate with the number of cases it was measured on, and the failed cases. NIST's 2024 generative AI profile advises against extrapolating performance from narrow, anecdotal assessments.
When should a company start writing evals?
A company should start writing evals before the first release of an AI feature, and where possible before the feature is built. Anthropic's engineering post says evals get harder to build the longer you wait, and that 20 to 50 simple tasks drawn from real failures is a great start. Both statements are that company's advice from its own work. Writing the expectation first also makes the team agree on what the feature is for.
Can a public benchmark score replace my own evals?
No. A public benchmark score compares AI models on a shared set of questions, and it contains nothing about your products, policies or customers. Your own evals test the product on your own cases. A model with a high benchmark score can still state your return policy wrongly. A chatbot stating a company policy wrongly was the failure in the Air Canada tribunal decision of February 2024.
What is the first thing to ask an AI vendor about testing?
Ask the vendor how many test cases the pass rate was measured on and where those cases came from. The count matters because a 92% pass rate on 50 cases has a 95% range of 84.5% to 99.5%, and the same rate on 1,000 cases has a range of 90.3% to 93.7%, using the confidence interval method described by Anthropic. Then ask to read the failed cases.
Are AI evals only for large companies?
No. AI evals suit any company whose AI product gives answers to customers, whatever its size. A first suite can be small: Anthropic's engineering post says 20 to 50 simple tasks drawn from real failures is a great start. The first task at that size is writing: people who know the business state what a correct answer must do for each case.
What happens if a company releases an AI feature without evals?
A company that releases an AI feature without evals learns about wrong answers from customer complaints, and has no record of what it checked. Regulators have also acted on advertising claims about AI products. The US Federal Trade Commission's final order of February 11, 2025, a settlement, requires DoNotPay to pay $193,000 and prohibits it from advertising that its service performs like a real lawyer unless it has sufficient evidence.
Do evals stop once the product is released?
No. Evals continue after release, because the product and the model it uses keep changing. The NIST AI Risk Management Framework says AI systems should be tested before their deployment and regularly while in operation. Run the suite again after every change to the product, whenever the vendor replaces or retires the model, and each time a customer reports a wrong answer, which becomes a new case.
Can a non-engineer take part in writing evals?
Yes. A non-engineer writes the part of an eval that the rest depends on, the expectation: one or two sentences stating what a correct answer must do for a given input. That knowledge comes from the business, such as a return policy or a pricing rule. Engineers then turn each expectation into a check a computer can run. NIST's framework also asks for people who did not build the system to take part in assessments.
Do evals make an AI product free of errors?
No. Evals measure how often a product passes the cases in the suite, and a situation with no case stays untested. In a 2024 study of five language models by Berk Atil and co-authors, none consistently delivered repeatable accuracy across all tasks. Evals show the error rate on known situations, so that a company can set a pass line, block specific failures and decide a release with evidence.
How does Reveneau use evals in its own work?
Reveneau writes all of its code with AI, and every change must pass a large eval suite, written from the specification before the code, before release. The checks that need judgement are graded by Jev, TypeSafe AI's decision model, and on Reveneau's own suite the run is ten times faster than with its previous language-model grader. Reveneau as a company takes responsibility for the project through production and after release.
References
- [1] British Columbia Civil Resolution Tribunal, Moffatt v. Air Canada, 2024 BCCRT 149 (14 February 2024): the finding that Air Canada did not take reasonable care to ensure its chatbot was accurate; the order of $812.02 and its three parts.
- [2] OpenAI, Evaluation best practices (API documentation, read 30 September 2026): models sometimes produce different output from the same input, which makes traditional software testing methods insufficient.
- [3] Atil and 12 other authors, Non-Determinism of "Deterministic" LLM Settings (arXiv, 6 August 2024, revised 2 April 2025): five LLMs, eight tasks, 10 runs; accuracy variations up to 15% across runs.
- [4] NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 (July 2024): action MS-2.5-001 on anecdotal assessments; action GV-1.3-002 on minimum thresholds reviewed as part of deployment approval.
- [5] US Federal Trade Commission, FTC Finalizes Order with DoNotPay (11 February 2025): $193,000 in monetary relief; the service may not be advertised as performing like a real lawyer without sufficient evidence.
- [6] US Federal Trade Commission, FTC Order Requires Workado to Back Up Artificial Intelligence Detection Claims (28 April 2025): a product promoted as 98 percent accurate; independent testing showed 53 percent on general-purpose content, according to the FTC's complaint.
- [7] Chen, Zaharia and Zou, How is ChatGPT's behavior changing over time? (arXiv, 18 July 2023, revised 31 October 2023): on 1,000 prime-number questions GPT-4 went from 84.0% to 51.1% and GPT-3.5 from 49.6% to 76.2% between March and June 2023.
- [8] Anthropic, Model deprecations (Claude Platform Docs, read 30 September 2026): at least 60 days' notice before retirement of publicly released models; requests to retired models fail.
- [9] Anthropic, A statistical approach to model evaluations (19 November 2024): a 95% confidence interval is the mean score plus and minus 1.96 times the standard error of the mean.
- [10] Buolamwini and Gebru, Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification (Proceedings of Machine Learning Research, volume 81, 2018): 3 commercial systems; error rates up to 34.7% for darker-skinned females and a maximum of 0.8% for lighter-skinned males.
- [11] NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (January 2023): MANAGE 1.1 on the decision to proceed; the framework does not prescribe risk tolerance; MEASURE 1.3 on assessors who did not build the system; testing before deployment and regularly while in operation.
- [12] Kapoor, Stroebl and 29 other authors, Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation (arXiv, 13 October 2025): 21,730 agent rollouts across 9 models and 9 benchmarks at a total cost the authors round to $40,000.
- [13] Anthropic, Demystifying evals for AI agents (9 January 2026): 20-50 simple tasks drawn from real failures as a start; evals get harder to build the longer you wait.
- [14] European Union, Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 9(8), as shown on the European Commission's AI Act Service Desk (read 30 September 2026): testing against prior defined metrics and probabilistic thresholds.
- [15] European Union, Regulation (EU) 2026/1744 (Digital Omnibus on AI), Official Journal of 24 July 2026, read on the EU Publications Office copy: high-risk rules apply from 2 December 2027 (Annex III systems) and 2 August 2028 (Annex I systems), replacing 2 August 2026.
Related reading
What the Air Canada chatbot ruling means for a company with an AI assistant
A British Columbia tribunal treated a chatbot's answer as the airline's own statement and ordered a payment of $812.02. This post sets out what the decision says, what it leaves out, and the written test that compares an assistant's answers with the policy text.
How to evaluate a vendor's eval suite
"We test everything" means nothing until you know who wrote the tests, what they check, and who is allowed to grade them.
AI evaluation and guardrails for production: how to know your AI actually works
"It seems to work" is not a production standard. Here is how we turn that feeling into a number, and how we stop the system from doing damage on the days the number is bad.
An AI demo is not a product
A convincing AI demo takes an afternoon. Turning it into something people trust in production takes most of the work, and most failures happen at that stage.
The real cost of shipping unverified code
The cost of unverified code does not arrive as a bug report. It arrives as a codebase nobody will touch, a review queue that never empties, and a team that has stopped trusting its own pipeline.
How to negotiate a software contract you can actually verify
Most build contracts describe effort, timeline, and payment, and leave the one hard question unanswered: on what evidence do you agree the thing is finished?