Choose how to score

RAG evaluation: test the search step and the answer separately

A RAG system (retrieval-augmented generation) has two steps: a search step fetches passages from your documents, and a language model writes an answer from them. Either step can cause a wrong answer, and the repair is different for each, so test them separately. Test the search step in code by checking that the expected passage came back and at which position. Score the answer for faithfulness and answer relevance, and the fetched passages for context relevance, as the 2023 Ragas paper defines them. On the paper's own test of 50 Wikipedia pages, those measures agreed with human annotators at 0.95, 0.78 and 0.70.

Published September 30, 2026. Editorial.

Key takeaways

  • A wrong answer from a RAG system has two possible causes: the search step fetched the wrong passages, or the model answered wrongly from the right ones. One overall quality score cannot tell them apart.
  • Test the search step in code with no model: for each test question, record the passage that holds the answer, then check whether the search fetched it and at which position.
  • Faithfulness, as the 2023 Ragas paper defines it, is the number of statements in the answer that the fetched passages support, divided by the total number of statements in the answer.
  • On the Ragas authors' own test of 50 Wikipedia pages with two annotators, agreement with the annotators was 0.95 for faithfulness, 0.78 for answer relevance and 0.70 for context relevance.
  • Every measure here compares the answer with your documents. An answer that correctly repeats an out-of-date document passes all of them, so the documents need their own review.

Take a support assistant for an invented furniture shop. A customer asks how many days they have to return a sofa, and the assistant answers "30 days". The shop's returns policy says 14. This is an invented illustration, and it has two possible causes. The search step may have fetched the wrong document, for example an old policy page or a page about delivery. Or the search step fetched the right page and the model wrote 30 anyway.

The two causes need different repairs. The first is a search problem: how documents are cut into passages and how many passages are fetched. The second is a problem in the instructions given to the model. One overall score for "answer quality" cannot say which repair to make. This page, part of the guide to LLM evals for AI products, sets out how to test each step alone, using the measures defined in the Ragas paper [2] and the ARES paper [4].

What RAG means and where the name comes from

RAG stands for retrieval-augmented generation: a search step fetches passages from your documents, and a large language model (LLM) writes its answer from those passages. The name comes from a 2020 paper by Patrick Lewis and eleven co-authors, accepted at the NeurIPS 2020 conference [1]. Their system paired a language model with a searchable copy of Wikipedia from December 2018, cut into pieces of 100 words, 21 million pieces in total [1].

That paper also reported a comparison judged by people. On 452 pairs of generated quiz questions, evaluators judged the RAG output more factual than the output of BART, a model with no search step, in 42.7 percent of cases. They judged BART more factual in 7.1 percent [1]. The 2020 system trained the search component and the model together. A product that puts a search step in front of a model run by its provider does no such training, so those figures describe that research system and you have to measure your own.

Two words appear in every measure below. "Retrieval" is the search step. "Context" is the set of passages handed to the model.

The search step and the answer step fail in different ways

A user sees one wrong answer, and the team has to find out which step produced it. This table maps what the user sees to the step that failed and the measure that finds it.

What the user sees Step that failed Measure that finds it
The answer states something the documents do not say Answer step Faithfulness [2]
The answer agrees with the documents and replies to a different question Answer step Answer relevance [2]
"I could not find that", when the document exists Search step Expected-passage check; context recall [3]
The answer mixes the right fact with unrelated material Search step Context relevance [2]; context precision [3]
The answer correctly repeats an out-of-date document The document store A review of the documents themselves

The last row is a fault in the documents. Every measure on this page compares the answer with the documents, and the documents themselves go unchecked. A system can pass all of them and still tell a customer the wrong return period, because the stored page says so.

How to test the search step

The search step can be tested without running the model, and this is the cheapest test on the page. For each test question, a person who knows the documents records which passage holds the answer. A check written in code then runs the search and records whether that passage was fetched and at which position in the list. No model takes part, so the result is the same on every run.

Here is an invented illustration with 50 test questions and a search step that fetches 5 passages per question.

Check Count Share
Expected passage among the 5 fetched 41 of 50 41 / 50 = 82 percent
Expected passage in first position 27 of 50 27 / 50 = 54 percent
Expected passage missing 9 of 50 9 / 50 = 18 percent

The 9 misses are the place to start reading. A model cannot answer from a passage it never received, so those 9 cases can only be repaired in the search step. Reading them one by one is the method in error analysis before metrics. The test set should also hold questions the documents cannot answer, where the correct behaviour is to say so. Where real questions come from is covered in how to build an LLM eval dataset from real usage.

The Ragas documentation, as read on 30 September 2026, describes two measures for this step [3]. Context precision "evaluates the retriever's ability to rank relevant chunks higher than irrelevant ones for a given query in the retrieved context". Context recall "measures how many of the relevant documents (or pieces of information) were successfully retrieved". In those sentences the retriever is the search step, a chunk is one passage cut from a longer document, and the query is the user's question.

The documentation's formula for context recall is the number of claims in a reference answer that the fetched passages support, divided by the total number of claims in that answer. It adds that "calculating context recall always requires a reference to compare against" [3]. Someone has to write the correct answer for each test question before this measure can run.

How to test the answer with the three measures in the Ragas paper

Ragas is a free, publicly available code package for scoring RAG systems. It was described in a 2023 paper by Shahul Es, Jithin James, Luis Espinosa-Anke and Steven Schockaert, presented at the EACL 2024 conference [2]. The paper names three measures. All three work "without having to rely on ground truth human annotations" [2], which means no person has to write a correct answer first. Each uses a language model as the grader.

Faithfulness

The paper's definition is that "the answer should be grounded in the given context" [2]. A model splits the answer into separate statements. A model then checks each statement against the fetched passages. The score is the number of supported statements divided by the total number of statements.

As an illustration, an answer with 5 statements, 4 of them supported, scores 4 / 5 = 0.8. In the furniture example, "you have 30 days" is the unsupported statement. Faithfulness is the measure that finds a hallucination, which is a statement the model made up.

Answer relevance

The paper's definition is that "the generated answer should address the actual question that was provided" [2]. A model reads only the answer and writes several questions that the answer could be a reply to. Each generated question is compared with the user's real question, and the score is the average similarity.

The comparison turns each question into a list of numbers that stands for its meaning (an embedding) and measures how closely two lists point in the same direction (cosine similarity). An answer that covers other subjects, or leaves half the question unanswered, produces generated questions that differ from the real one. The paper states that this measure "does not take into account factuality" [2], so a wrong answer to the right question can score well on it.

Context relevance

The paper's definition is that "the retrieved context should be focused, containing as little irrelevant information as possible" [2]. A model extracts the sentences in the fetched passages that are needed to answer the question. The score is the number of extracted sentences divided by the total number of sentences in the context.

As an illustration, if the search step fetched 12 sentences and 3 were needed, the score is 3 / 12 = 0.25. This measure scores the search step.

How closely the Ragas measures matched human judgment

The Ragas authors tested their measures on a set of cases they built, WikiEval, made from 50 Wikipedia pages and labelled by two people [2]. For each measure they counted how often it preferred the same one of two answers as those two annotators did. The agreement was 0.95 for faithfulness, 0.78 for answer relevance and 0.70 for context relevance [2]. Asking a model directly for a score from 0 to 10, the simpler method the paper calls "GPT Score", reached 0.72, 0.52 and 0.63 [2]. The authors write: "We found context relevance to be the hardest quality dimension to evaluate."

Read those figures with three facts beside them. The test used 50 pages and two annotators. The authors of the paper are the makers of the tool. The models were gpt-3.5-turbo-16k and text-embedding-ada-002 [2].

The figures show that the method can match people on one set of cases. Whether it matches people on your documents is a separate measurement, described in human review and agreement between reviewers.

What the Ragas documentation lists as of 30 September 2026

The code package has changed since the paper. On 30 September 2026 the documentation's list for retrieval-augmented generation included Context Precision, Context Recall, Response Relevancy and Faithfulness [3]. "Context relevance", the third measure in the 2023 paper, is absent from that list. The paper's "answer relevance" appears under the page title "Response Relevancy", and the text of that page calls it "Answer Relevancy" [3].

The code has moved too. On 30 September 2026 the public code store was under the account vibrantlabsai on GitHub, a code hosting site, the older address redirected there, and the licence file was the Apache License 2.0 [5].

The instruction that follows: write the package version and the exact measure names into every eval report. A "context relevance" score from 2023 and a "context precision" score from 2026 are two different calculations, and a chart that joins them shows a change that never happened.

ARES: a small grader corrected with human labels

ARES is a second method, from a 2023 paper by Jon Saad-Falcon, Omar Khattab, Christopher Potts and Matei Zaharia, published at the NAACL 2024 conference [4]. It scores the same three things, in the paper's words "context relevance, answer faithfulness, and answer relevance".

The method has three stages [4]. A model writes artificial question and answer pairs from your own documents. Small grader models are trained on those pairs. Then the graders' scores are corrected with a set of cases labelled by people, using a statistical method named prediction-powered inference. That correction also produces a confidence interval, which is the range the true score plausibly lies in. The paper asks for a labelled set of 150 cases or more, a figure it gives as approximate [4].

The idea worth taking from ARES is the correction step: a small set of human labels tells you how far to trust a cheap automatic grader.

What these measures cannot tell you

Three limits apply to every score on this page.

  • A faithful answer can be wrong. Faithfulness checks the answer against the fetched passages, so an answer that repeats an outdated page scores 1.
  • Every answer measure here is a model grading a model. A grader of that kind has to be checked against labels from people before its score decides anything. The general case is in never let the model grade its own work, and the design of the grader is in LLM as judge vs a decision model.
  • A score from 50 questions has a wide range of plausible true values. The arithmetic is in how many test cases an LLM eval needs.

For the engineering work behind the search step, see RAG enterprise search development.

How Reveneau applies this

At Reveneau all code is written by AI, and every change must pass a large eval suite that is written from the specification before the code exists. For a product with a search step, we write that suite in two parts. The first part tests the search step alone, in code: each test question names the passage that should come back. The second part tests the answer against the passages that were fetched: whether each statement is supported, and whether the answer replies to the question asked.

We grade the judged checks with Jev, TypeSafe AI's decision model. On our own suite, the run is ten times faster than it was with our previous language-model grader.

Reveneau, as a company, takes responsibility for the whole project through production and after release. For a RAG system that means a wrong answer found after release becomes a new test question with its expected passage, before the repair is written. To plan this for your own documents, see AI development at Reveneau or contact us.

Best for

  • Products where a model answers from your own documents: support assistants, document search, policy questions
  • Teams that need to know whether to repair the search step or the instructions to the model
  • Test sets where a person can name the passage that holds each answer

Avoid if

  • The product has no search step: use the general LLM eval methods instead
  • Nobody has labelled a sample by hand yet, so no model grader can be checked
  • The stored documents are known to be out of date: review those first

Check before you decide

  • Each test question records the expected passage
  • Search checks and answer checks are reported as separate numbers
  • The report names the library version and the exact measure names
  • The model grader was compared with human labels on your own cases

Common questions

What is RAG evaluation?

RAG evaluation is the testing of a system in which a search step fetches passages from your documents and a language model answers from them. The search step and the answer are tested separately, because either one can cause a wrong answer. The name RAG, retrieval-augmented generation, comes from a 2020 paper by Patrick Lewis and co-authors that paired a language model with a searchable copy of Wikipedia.

What does faithfulness mean when scoring a RAG answer?

Faithfulness measures whether the answer says only what the fetched passages support. The 2023 Ragas paper computes it in two steps: a model splits the answer into separate statements, then checks each one against the passages. The score is the supported statements divided by all statements, so an answer with 4 supported statements out of 5 scores 0.8.

How do I test the retrieval step of a RAG system?

Test the retrieval step in code, with no model involved. For every test question, a person who knows the documents records which passage holds the answer. The check runs the search and records whether that passage came back and at which position. In this page's invented illustration, the expected passage was fetched for 41 of 50 questions, which is 82 percent, and the 9 misses are read first.

What is Ragas and who maintains it?

Ragas is a free, publicly available code package for scoring RAG systems, first described in a 2023 paper by Shahul Es, Jithin James, Luis Espinosa-Anke and Steven Schockaert. As read on 30 September 2026, its public code store is under the GitHub account vibrantlabsai, the older address redirects there, and its licence file is the Apache License 2.0. The list of measures in its documentation has changed since the paper.

Why does my RAG system give wrong answers from the right documents?

A wrong answer from the right documents means the answer step failed while the search step worked. The model added a statement the passages do not support, or it replied to a different question. Faithfulness finds the first problem and answer relevance finds the second, as the Ragas paper defines them. The repair is in the instructions to the model, and changing the search step will leave the fault in place.

What are context precision and context recall?

Context precision and context recall are two measures of the search step in the Ragas documentation, as read on 30 September 2026. Context precision scores whether relevant passages are ranked above irrelevant ones. Context recall is the number of claims in a reference answer that the fetched passages support, divided by all claims in that reference answer. The documentation says context recall always needs a reference answer.

How well do the Ragas measures agree with human reviewers?

The Ragas authors report agreement with human annotators of 0.95 for faithfulness, 0.78 for answer relevance and 0.70 for context relevance. Those figures come from their own set of 50 Wikipedia pages labelled by two people, in a paper from 2023. The authors made the tool they tested, so the figures show the method can work and leave open how it performs on your documents.

What is the difference between Ragas and ARES?

Ragas, as its 2023 paper describes it, asks a language model to score each answer directly, with no human labels. ARES, from a 2023 paper by Jon Saad-Falcon and co-authors, trains small grader models on artificial question and answer pairs and then corrects their scores with 150 or more cases labelled by people. That correction also gives a range for the true score.

Do I need reference answers written by a person to evaluate a RAG system?

Reference answers are needed for some measures and optional for others. The three measures in the Ragas paper, faithfulness, answer relevance and context relevance, work without one. Context recall, as the Ragas documentation defines it, always requires a reference answer. A code check of the search step needs something smaller: the passage that holds the answer, recorded once for each test question.

Can a RAG answer pass every measure and still be wrong?

Yes, a RAG answer can pass every measure on this page and still be wrong. Each measure compares the answer with the fetched passages, and none compares the passages with the facts outside the documents. An answer that correctly repeats an out-of-date policy page scores 1 on faithfulness. Keeping the stored documents current is a separate job, with its own owner and its own review.

How much work is it to set up a RAG eval?

The main work in a RAG eval is labelling: a person who knows the documents records the passage that holds the answer for each test question. The search check then runs in code for the cost of computer time only. The answer measures need a model call for each answer, and a sample of answers labelled by people to confirm the model grader agrees with them. ARES asks for 150 or more such labels.

What should a team do first when its RAG answers are poor?

Start with the search step. Record the expected passage for a set of real questions and check how often the search fetches it. A model cannot answer from a passage it never received, so those misses must be repaired in the search step. Then score the remaining answers for faithfulness and read the unsupported statements one by one before changing the instructions to the model.

References