LLM evals: how to measure whether an AI product works / Choose how to score
LLM eval metrics: exact match, similarity scores, rubrics and pairwise comparison
Metrics for evaluating a product built on a large language model (LLM) belong to five families. Exact match and code assertions check an output that has one right form. BLEU and ROUGE count the words an output shares with a reference text written by a person. BERTScore compares meaning instead of exact words. A rubric score is a judgement by a person or a model against written criteria. Pairwise comparison asks which of two outputs is better. Each family answers a different question, so choose one metric for each failure type you found when reading real outputs, and report each one as its own pass rate. A single overall score hides which failure type changed.
Published September 30, 2026. Editorial.
Key takeaways
- Exact match and code assertions check outputs that have one right form. Anthropic's documentation calls code-based grading the fastest and most reliable method.
- BLEU and ROUGE count words shared with a reference text. Ehud Reiter's 2018 review of 284 correlations in 34 papers found the evidence does not support BLEU outside machine translation or for individual texts.
- In the G-Eval paper, a GPT-4 rubric grader reached an average Spearman correlation of 0.514 with human ratings on summaries, where 1 would mean identical rankings, so a model's rubric score needs checking against human labels.
- Judge every pairwise comparison in both orders. Zheng and co-authors found that GPT-4 kept its verdict in 65.0 percent of cases when the two answers swapped places.
- Choose one metric for each failure type found in error analysis and report each as its own pass rate. One overall score hides which failure type changed.
Take a support assistant for an invented furniture shop. A person has written the reference answer to one test question: "The refund takes five days." On two runs the assistant produces two outputs. The first is "The refund takes ten days." The second is "You get your money back within five days." A score that counts shared words, called ROUGE-1 recall, gives the first output 0.8, because four of the five reference words appear in it. It gives the second output 0.4, because only "five" and "days" appear. The wrong answer scores twice as high as the right one.
That is an invented illustration, computed by hand from the definition in the ROUGE paper [3]. A metric is the rule that turns an output into a score, and each metric counts one specific thing. This page describes five families of metric used in evals, the repeatable tests of an AI product's output. It is one step of the method in the guide to LLM evals, where LLM stands for large language model.
Five families of LLM eval metric in one table
| Metric | What it counts | Good for | Fails when |
|---|---|---|---|
| Exact match and code assertions | Whether the output equals a known answer or obeys a rule written in code | Labels, amounts, formats, required links | A correct answer can be worded many ways |
| BLEU | Runs of words the output shares with reference translations | Comparing translation systems over a whole test set | Used on single texts or outside translation [2] |
| ROUGE | Words from a reference summary that appear in the output | Summaries with reference summaries written by people | The output is right in new words, or wrong in familiar ones |
| BERTScore | Similarity of meaning between the output and a reference | Outputs that reword the reference | No reference exists, or nobody has tested the score on your cases |
| Rubric score | A judgement against written criteria, by a person or a model | Tone, completeness, following instructions | The grader was never compared with human labels |
| Pairwise comparison | Which of two outputs a grader prefers | Choosing between two versions | You need to know whether either one is good enough |
Exact match, BLEU, ROUGE and BERTScore need a reference: a correct answer written or confirmed by a person. Google Cloud's documentation says its computed metrics, "like ROUGE or BLEU", apply "when a ground truth is available" [10]. Ground truth means an answer accepted as correct. OpenAI's documentation names exact match and ROUGE or BLEU scoring among its examples of metric-based evals, and lists human evals and model graders as the other two types [8]. A grader is the person or model that gives the score.
Exact match and code assertions
Anthropic's documentation defines the first: "Exact match evals measure whether the model's output matches a predefined correct answer, typically after normalizing whitespace and case." [7] Normalizing here means ignoring differences in spacing and capital letters. A code assertion is a rule written in code that the output must satisfy. The answer contains the returns link. The answer stays under 120 words. The answer states a return period of 30 days.
The same documentation compares three ways to grade. It calls code-based grading the fastest and most reliable, human grading the most flexible but "slow and expensive", and grading by a model fast and flexible once its reliability has been tested [7]. Its rule is to "choose the fastest, most reliable, most scalable method" that can decide the question [7].
Use a code check for every failure type that has one right form. The page on the three kinds of LLM eval compares it with the other two.
BLEU and ROUGE count shared words
BLEU was published in 2002 by four researchers at IBM for machine translation [1]. BLEU counts the runs of one to four words in the output that also appear in reference translations, combines the four shares into one figure, and applies a penalty to outputs that are too short. The result lies between 0 and 1 [1].
The BLEU paper states two limits. Scores "on individual sentences will often vary from human judgments", because BLEU was built to be averaged over a whole test set [1]. And the score depends on how many references exist: on a test set the paper describes as 500 sentences, a human translator scored 0.3468 against four references and 0.2571 against two [1].
Ehud Reiter reviewed the later evidence in 2018, covering "284 correlations reported in 34 papers". The abstract of that review says the evidence supports BLEU for diagnostic evaluation of machine translation (MT) systems and "does not support using BLEU outside of MT, for evaluation of individual texts, or for scientific hypothesis testing" [2].
ROUGE was published by Chin-Yew Lin in 2004 for summaries: "ROUGE stands for Recall-Oriented Understudy for Gisting Evaluation." [3] Its measures count the units that a computer-written summary shares with "ideal summaries created by humans" [3]. ROUGE-N is the share of the reference's word runs that appear in the output. ROUGE-L uses the longest sequence of words that both texts contain in the same order [3]. The opening illustration is ROUGE-N with runs of one word.
A product answer is a single text, usually outside translation. OpenAI's guide lists "Overly generic metrics" among common mistakes and names BLEU as an example [8].
BERTScore compares meaning instead of exact words
BERTScore was published in 2019 by Tianyi Zhang and four co-authors, who set out to credit an output that keeps the meaning in different words [4]. Their own example uses the reference "people like foreign cars". BLEU, and a second word-overlap score called METEOR, give a higher score to "people like visiting places abroad" than to "consumers prefer imported cars" [4].
BERTScore changes what counts as a match. A language model named BERT turns each token, meaning a word or a piece of a word, into a list of numbers called an embedding. Tokens used with similar meaning get similar numbers. In the authors' words: "instead of exact matches, we compute token similarity using contextual embeddings" [4].
The authors tested it on "the outputs of 363 machine translation and image captioning systems" [4]. They state two limits. First, "there is no one configuration of BERTScore that clearly outperforms all others" [4]. Second, raw scores are in a narrow range, which "makes the actual score less readable", so the authors rescale them against a baseline [4].
The published tests cover translation and image captions. Before a similarity score decides anything in your product, run it on cases a person has already labelled and count how often it agrees.
Rubric scores from a person or a model
A rubric is a written list of criteria with a description of what each score means. Google Cloud's documentation for its own evaluation service describes two kinds. It describes static rubrics in these words: "Apply a fixed set of scoring criteria across all prompts." They return one number for each prompt, such as a score from 1 to 5. Adaptive rubrics, which Google recommends, generate "a unique set of pass or fail rubrics for each individual prompt" [10]. A prompt here is one test input.
The G-Eval paper of 2023 tested a model as the grader. Its grader prompt holds the criteria, evaluation steps that the model writes for itself, and a form to fill in [5]. On SummEval, a public test set of summaries, G-Eval with GPT-4 reached an average Spearman correlation of 0.514 with human ratings, the best figure in the paper's table [5]. A Spearman correlation measures how closely two rankings agree: 1 means the same order and 0 means no relation.
The same paper reports two problems. "LLMs usually only output integer scores, even when the prompt explicitly requests decimal values." [5] And the GPT-4 grader gave GPT-3.5 summaries higher scores than summaries written by people, "even when human judges prefer human-written summaries" [5]. Both models were the 2023 versions.
A model's rubric score therefore has to be compared with human labels before it decides a release. The method is on the page about human review and reviewer agreement. The sample code in Anthropic's documentation carries a second precaution in a comment: "Generally best practice to use a different model to evaluate than the model used to generate the evaluated output" [7]. The post on never letting the model grade its own work argues the same for code.
Pass or fail is easier to keep consistent than a 1 to 5 scale
Hamel Husain and Shreya Shankar take a clear position: "Binary evaluations force clearer thinking and more consistent labeling." [9] Binary means two outcomes, pass or fail. They name three problems with a 1 to 5 scale: "the difference between adjacent points (like 3 vs 4) is subjective and inconsistent across annotators, detecting statistical differences requires larger sample sizes, and annotators often default to middle values to avoid making hard decisions" [9]. An annotator is a person who labels outputs.
OpenAI's guide gives similar advice. For human review it says: "Include a pass/fail threshold in addition to the numerical score". For model graders it says: "Use pairwise comparison or pass/fail for more reliability" [8]. These are positions held by two practitioners and one vendor. The sources read for this guide contain no controlled study that compares the two scales.
We take the same position. Write each criterion as its own pass or fail question. A 1 to 5 score for "quality" becomes several questions. Is every figure in the answer present in the order record? Does the answer give the returns link? Each result names the thing to fix. A pass or fail result also produces a rate with a range around it, which the page on how many test cases an eval needs computes.
Pairwise comparison asks which of two outputs is better
Lianmin Zheng and co-authors describe the method in a 2023 paper: "An LLM judge is presented with a question and two answers, and tasked to determine which one is better or declare a tie." [6]
The paper measured two biases in the graders of 2023.
- Position bias is favouring an answer because of where it appears. When the two answers swapped places, GPT-4 kept its verdict in 65.0 percent of cases, GPT-3.5 in 46.2 percent and Claude-v1 in 23.8 percent [6]. The authors' remedy: "call a judge twice by swapping the order of two answers and only declare a win when an answer is preferred in both orders" [6].
- Verbosity bias is favouring longer answers. The authors lengthened 23 answers with a repeated list and counted how often each grader failed the test: 91.3 percent for Claude-v1, 91.3 percent for GPT-3.5 and 8.7 percent for GPT-4 [6].
Pairwise comparison has two further limits. The paper notes that the number of pairs grows with the square of the number of versions [6]. By our arithmetic, 4 versions make 6 pairs and 10 versions make 45. And a pairwise result names a winner. Whether the winner is good enough to release is a separate pass or fail check.
Choose one metric for each failure type
Error analysis, reading real outputs and grouping the failures, produces a list of failure types. Each type gets the cheapest metric that can decide it. An invented list for the furniture shop:
- Wrong return period stated: a code assertion on the number of days.
- Returns link missing: a code assertion that the link is present.
- Rude tone: a pass or fail rubric question, graded by a model that has been compared with a person's labels.
- Two candidate greetings: a pairwise comparison, judged in both orders.
One overall score hides these types. Suppose two versions each pass 90 of 100 cases. Version A has 2 wrong return periods and 8 tone failures. Version B has 8 wrong return periods and 2 tone failures. The overall score is 90 percent for both, and version B states four times as many wrong facts. Report each failure type as its own pass rate.
The page on error analysis covers how to find the types. Products that search documents before answering need measures for the search step as well, described under RAG evaluation. For the design of a model grader, see LLM as judge vs a decision model.
How Reveneau chooses metrics
At Reveneau we write the expectation for each case before the prompt, and we give every failure type its own check. All of our code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before release. Inside that suite, anything with one right form is a check in code. The checks that need judgement are written as pass or fail questions and graded by Jev, TypeSafe AI's decision model. On our own suite, that run is ten times faster than with our previous language-model grader.
We report each failure type as its own pass rate, so a client sees which type changed. Using AI instead of adding engineers keeps the team small, and that saving goes into the client's price. Reveneau, as a company, takes responsibility for the whole project through production and after release. To choose metrics for your own product, see AI development at Reveneau.
Best for
- Teams that have a list of failure types and need a scoring method for each
- Deciding whether a ready-made score such as ROUGE fits a product
- Choosing between a 1 to 5 scale and pass or fail questions
Avoid if
- You have yet to read real outputs: do error analysis before choosing any metric
- You want one number to describe overall quality
Check before you decide
- Every failure type has its own check and its own pass rate
- Each model-graded check was compared with human labels
- Pairwise comparisons were judged in both orders
Common questions
What metrics are used to evaluate LLMs?
Five families of metric are used to evaluate LLM output: exact match and code assertions, word-overlap scores such as BLEU and ROUGE, meaning-similarity scores such as BERTScore, rubric scores given by a person or a model, and pairwise comparison between two outputs. OpenAI's documentation names exact match and ROUGE or BLEU scoring as metric-based evals, and lists human evals and model graders as the other two types.
What is BLEU?
BLEU is a score from a 2002 IBM paper that measures how close a machine translation is to reference translations written by people. BLEU counts the runs of one to four words that the output shares with the references and applies a penalty to short outputs, giving a figure from 0 to 1. The paper itself says scores on individual sentences will often vary from human judgement.
What is ROUGE?
ROUGE is a set of scores published by Chin-Yew Lin in 2004 for judging computer-written summaries against summaries written by people. The name stands for Recall-Oriented Understudy for Gisting Evaluation. ROUGE-N is the share of the reference's word runs that appear in the output, and ROUGE-L uses the longest sequence of words that both texts contain in the same order.
What is BERTScore?
BERTScore is a similarity score from a 2019 paper that compares an output with a reference by meaning. A language model turns each word or word piece into a list of numbers, and BERTScore matches the pieces of the two texts by how close those numbers are. The authors tested BERTScore on the outputs of 363 translation and image captioning systems and say no single configuration clearly outperforms all others.
Is a pass or fail score better than a 1 to 5 scale?
Pass or fail is the better default for product evals, according to Hamel Husain and Shreya Shankar, who write that binary evaluations force clearer thinking and more consistent labelling. They name three problems with a 1 to 5 scale: neighbouring points mean different things to different reviewers, detecting a difference needs a larger sample, and reviewers default to the middle values. OpenAI's guide also recommends a pass or fail threshold.
What is pairwise comparison in LLM evaluation?
Pairwise comparison shows a grader one question and two answers and asks which answer is better, with a tie allowed. Zheng and co-authors measured a weakness in the 2023 models: when the two answers swapped places, GPT-4 kept its verdict in 65.0 percent of cases and Claude-v1 in 23.8 percent. Their remedy is to judge each pair in both orders and count a win only when both agree.
Which LLM eval metric should I start with?
Start with code assertions for every failure type that has one right form, such as a wrong amount or a missing link. Anthropic's documentation calls code-based grading the fastest and most reliable method. Add a pass or fail rubric question, graded by a person first, for each failure type that needs judgement. Leave similarity scores until you have reference answers and labelled cases to test them on.
Why does a word-overlap score fail for a support assistant?
A word-overlap score counts shared words, so a support answer with a wrong fact in familiar wording can outscore a correct answer in new wording. In this page's invented illustration, the reference is 'The refund takes five days', and an answer saying ten days scores 0.8 on ROUGE-1 recall while a correct reworded answer scores 0.4. Ehud Reiter's 2018 review found the evidence does not support BLEU for individual texts.
Can I report one overall quality score for an AI product?
One overall quality score hides which failure type changed. In this page's invented illustration, two versions both pass 90 of 100 cases, yet one has 2 wrong facts and 8 tone failures and the other has 8 wrong facts and 2 tone failures. Report a separate pass rate for each failure type found in error analysis, and show the overall figure only next to those.
Do I need reference answers to use these metrics?
Reference answers are needed for exact match, BLEU, ROUGE and BERTScore, which all compare the output with a text a person wrote or confirmed. Google Cloud's documentation says its computed metrics apply when a ground truth is available, meaning an answer accepted as correct. Code assertions, rubric scores and pairwise comparison can run without a reference answer, because each one judges the output against a rule, written criteria or a second output.
How well does a model's rubric score agree with human ratings?
A model's rubric score agreed with human ratings only in part in the G-Eval paper. That paper reports an average Spearman correlation of 0.514 between its GPT-4 grader and human ratings on the SummEval summaries, where 1 would mean identical rankings. The same grader scored GPT-3.5 summaries above human-written ones even when human judges preferred the human text. Compare any model grader with your own human labels.
How much does each type of metric cost to run?
The cost of a metric follows who or what does the scoring. Anthropic's documentation describes code-based grading as the fastest and most reliable method, human grading as slow and expensive, and grading by a model as fast and flexible once its reliability has been tested. Word-overlap and similarity scores also need reference answers, which a person has to write. The sources give no prices, so measure cost on your own cases.
References
- [1] Kishore Papineni, Salim Roukos, Todd Ward and Wei-Jing Zhu, BLEU: a Method for Automatic Evaluation of Machine Translation (ACL, July 2002): what BLEU counts, its 0 to 1 range, the limit on single sentences, and 0.3468 against four references and 0.2571 against two on 500 sentences.
- [2] Ehud Reiter, A Structured Review of the Validity of BLEU (Computational Linguistics, September 2018, abstract): 284 correlations in 34 papers; the evidence does not support BLEU outside machine translation, for individual texts, or for scientific hypothesis testing.
- [3] Chin-Yew Lin, ROUGE: A Package for Automatic Evaluation of Summaries (ACL workshop, July 2004): the name, what the measures count, and the definitions of ROUGE-N and ROUGE-L.
- [4] Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger and Yoav Artzi, BERTScore: Evaluating Text Generation with BERT (arXiv, April 2019, revised February 2020): the method, the "people like foreign cars" example, the 363 systems tested, and the two stated limits.
- [5] Yang Liu and five co-authors, G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment (arXiv, March 2023, revised May 2023): the grader design, Spearman correlation of 0.514 on SummEval, integer scores, and higher scores for GPT-3.5 summaries than for human-written ones.
- [6] Lianmin Zheng and twelve co-authors, Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv, June 2023, revised December 2023): pairwise comparison, position consistency of 65.0, 46.2 and 23.8 percent, the test on 23 lengthened answers, and the remedy of judging in both orders.
- [7] Anthropic, Define success criteria and build evaluations (Claude Platform documentation, undated, read 30 September 2026): the definition of exact match, the three grading methods, the rule for choosing among them, and the comment about using a different model as grader.
- [8] OpenAI, Evaluation best practices (API documentation, undated, read 30 September 2026): the three evaluator types, examples of metric-based evals, the pass or fail threshold, pairwise comparison or pass or fail for model graders, and the mistake of overly generic metrics.
- [9] Hamel Husain and Shreya Shankar, AI Evals: Everything You Need to Know (page dated 18 September 2026): binary evaluations and the three stated problems with 1 to 5 scales.
- [10] Google Cloud, Gen AI evaluation service overview (read 30 September 2026): static and adaptive rubrics, and computation-based metrics for cases where a ground truth is available.
Related reading
Never let the model grade its own work
When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.
How to test a feature you cannot fully specify
Some features have no single correct output. A summary, a ranking, a suggested reply. You cannot check those against one exact expected answer, and the usual conclusion, that they cannot be tested, is wrong.
The real price of an LLM judge is per decision, and it multiplies
A grader's price per call looks small. Multiply it by every check on every change, or every message through a guardrail, and the public figures from LangChain, Openlayer and TypeSafe show where a cheaper decision model pays for its own migration and where it does not.
What to do with a grader that gives no reason
A Jev grade is a probability with no paragraph. That is enough for a check you wrote and calibrated yourself, and wrong for a new kind of failure, a design judgment, or anything a regulator wants explained in words.
More in Choose how to score
RAG evaluation: test the search step and the answer separately
A RAG system (retrieval-augmented generation) has two steps: a search step fetches passages from your documents, and a language model writes an answer from them. Either step can cause a wrong answer, and the repair is different for each, so test them separately. Test the search step in code by checking that the expected passage came back and at which position. Score the answer for faithfulness and answer relevance, and the fetched passages for context relevance, as the 2023 Ragas paper defines them. On the paper's own test of 50 Wikipedia pages, those measures agreed with human annotators at 0.95, 0.78 and 0.70.
Human review in LLM evals: guidelines, labels and agreement between reviewers
Human labels are the reference that every model grader is checked against, so their quality sets the limit for every automatic score. Have a person who knows the field label outputs as pass or fail with a written reason, following a written guideline. Have a second person label the same sample separately, then measure agreement with Cohen's kappa, which removes the agreement expected by chance. In this page's invented example, two reviewers agree on 85 of 100 outputs and kappa is 0.571. The bands used to describe kappa, such as Landis and Koch's from 1977, are a convention, so set your own level by what a wrong label would cost.