Evaluation, Benchmarking & Observability · 7 min
Evaluating RAG: Measure Retrieval and Generation Apart
If you only score the final answer, you cannot tell a retrieval miss from a generation lie. Decompose the pipeline and each failure gets its own number.
A support bot tells a customer that your refund window is 90 days. Your policy says 30. The answer is confidently wrong, and now you have a ticket. Was the mistake that retrieval pulled the wrong policy page, maybe an outdated 2019 doc that still said 90? Or did retrieval pull the correct 30-day page and the model invented 90 anyway? Those are two different bugs with two different fixes, and a single "the answer was wrong" score tells you nothing about which one you have. This is the core problem with evaluating retrieval-augmented generation: the thing you can see easily, the final answer, sits downstream of at least two systems that fail independently.
Why one end-to-end number lies to you
RAG is a pipeline. A query goes to a retriever, which returns some chunks of context. Those chunks plus the query go to a generator, which writes an answer. If you only measure the answer (human thumbs-up, an exact-match score against a gold answer, a single LLM judge asked "is this good?") you have collapsed two stages into one scalar. That scalar behaves like a product, not a sum, and products hide their factors.
Consider the four ways a RAG answer can go wrong:
retrieval good + generation good -> correct answer retrieval good + generation bad -> hallucination (chunks were right, model ignored them) retrieval bad + generation good -> faithful to garbage (model dutifully used wrong context) retrieval bad + generation bad -> wrong for two reasons
An end-to-end score of "40% correct" is consistent with all sorts of underlying states. Maybe your retriever is excellent and your prompt is starving the model of instructions. Maybe your generator is fine and your chunking strategy splits every table in half so the numbers never travel with their labels. You cannot tell, so you cannot prioritize. You end up tuning the embedding model when the real fix was three lines in the prompt, or the reverse.
The decomposition is the whole game. Measure retrieval quality on its own axis, measure generation quality on its own axis, and each failure gets a number you can drive down.
Scoring the retriever: precision and recall over context
Retrieval is an information-retrieval problem, so it gets information-retrieval metrics, adapted to the fact that "relevant" now means "useful for answering this specific question."
Context recall asks: of everything needed to answer the question, how much did retrieval actually bring back? You need a reference answer (or a set of ground-truth facts) to compute it. Break the reference into claims, then check each claim against the retrieved chunks: can this claim be attributed to something we retrieved? Recall is the fraction that can.
question: "What's the refund window and who approves exceptions?" reference facts: 1. refund window is 30 days -> found in retrieved chunk 2 [yes] 2. exceptions approved by manager -> not in any retrieved chunk [no] context recall = 1/2 = 0.50
Low recall means the answer was never fully in the context to begin with. No amount of prompt engineering fixes that; the information wasn't in the room. You fix recall upstream: better chunking, a bigger top_k, hybrid keyword-plus-vector search, query rewriting, or a re-ranker.
Context precision asks the complementary question: of the chunks you retrieved, how many are actually relevant, and do the relevant ones sit near the top? Precision matters because junk context is not free. It burns tokens, and it gives the model room to anchor on something off-topic. RAGAS computes context precision by looking at where the relevant chunks land in the ranked list and rewarding relevant items that appear early, a rank-aware measure rather than a flat "what fraction were relevant."
If you have graded relevance labels for a benchmark set, the classic IR metrics apply directly and are worth computing alongside the LLM-judged ones: hit@k (did any relevant chunk make the top k), MRR (how high the first relevant chunk ranked, averaged over queries), and nDCG (rank-weighted relevance across the whole list). RAGAS's context precision is a close cousin of these. It just derives the relevance judgments from an LLM instead of a human-labeled qrels file, which is what lets you run it without a hand-built relevance set for every query.
The two trade off exactly as they do in classic IR. Crank top_k to 50 and recall climbs while precision craters, because you have dragged in more relevant chunks and far more junk at the same time. The point of measuring both is to see the trade instead of stumbling into it.
Scoring the generator: faithfulness and answer relevance
Now hold retrieval fixed and ask whether the generator did its job with the context it got.
Faithfulness (also called groundedness) measures whether the answer is supported by the retrieved context. This is the hallucination detector. The mechanism RAGAS uses is claim decomposition: break the generated answer into individual factual statements, then, for each one, ask whether it can be inferred from the retrieved context. Faithfulness is the fraction of claims that can.
answer: "The refund window is 30 days, and exceptions
are approved by a regional manager."
claim 1: refund window is 30 days -> supported by context [yes]
claim 2: exceptions approved by regional mgr -> NOT in context [no]
faithfulness = 1/2 = 0.50Notice what this does and does not tell you. A faithfulness of 0.50 says half the answer's claims are ungrounded: the model made them up or pulled them from parametric memory. And faithfulness is measured against the retrieved context, not against the truth. An answer can be perfectly faithful to context that is itself wrong (the "faithful to garbage" quadrant). That is a feature. It cleanly separates "the model invented things" from "the model was fed bad chunks," which is precisely the separation an end-to-end score destroys.
Answer relevance measures whether the answer actually addresses the question asked, penalizing responses that are incomplete or padded with irrelevant material. RAGAS estimates it with a trick: prompt an LLM to generate several questions that the answer would be a good response to, embed those synthetic questions, and measure their average cosine similarity to the original question. An answer that is on-point generates questions that look like the real one; a rambling or evasive answer generates questions that drift.
Faithfulness and answer relevance catch different failures. A model can be faithful but unresponsive ("According to the context, our office is open 9 to 5" when you asked about refunds), or responsive but unfaithful. You want both numbers, because a single "is this answer good?" judge blurs the two together and hands you back the same uninformative scalar you were trying to escape.
Putting it together, and where the seams show
RAGAS ("Ragas: Automated Evaluation of Retrieval Augmented Generation," Es et al., 2023) packaged these four metrics (context precision, context recall, faithfulness, answer relevance) into a reference-lite framework whose selling point is that most of them need no human-written gold answers. The claim extraction and the judgments are done by an LLM. That is what makes it cheap enough to run on every build.
A minimal run looks like this:
from ragas import evaluate
from ragas.metrics import (
faithfulness, answer_relevancy,
context_precision, context_recall,
)
# dataset rows: question, retrieved contexts, generated answer,
# and a ground_truth reference (needed for context_recall)
result = evaluate(
dataset,
metrics=[context_precision, context_recall,
faithfulness, answer_relevancy],
)
print(result)
# {'context_precision': 0.82, 'context_recall': 0.54,
# 'faithfulness': 0.91, 'answer_relevancy': 0.88}Read that output the way it is meant to be read. Retrieval recall is the weak link at 0.54, the context often does not contain the answer, while the generator behaves well at 0.91 faithfulness. Now you know to spend the week on chunking and search, not on the prompt. That is a decision the numbers made for you, and it is the whole payoff of decomposing.
Two honest caveats. First, these metrics are computed by an LLM judge, so they inherit the judge's blind spots: miscounting claims, rewarding fluent nonsense, drifting when you swap judge models. Treat the scores as directional signal, not gospel. Pin your judge model and version so week-to-week comparisons mean something, anchor a small human-labeled set, and periodically check that the judge still agrees with people. Second, claim decomposition is fuzzy. How you split an answer into claims changes the denominator, so absolute values matter less than the trend on a fixed dataset. Use these numbers to catch regressions and to localize which half broke. That is what they are genuinely good at. Do not oversell a 0.91 as if it were a physical constant.
Sources
- Es, Shahul; James, Jithin; Espinosa-Anke, Luis; Schockaert, Steven. "Ragas: Automated Evaluation of Retrieval Augmented Generation." 2023. https://arxiv.org/abs/2309.15217
- Ragas documentation. "Metrics" (faithfulness, answer relevance, context precision, context recall). https://docs.ragas.io/en/stable/concepts/metrics/
- Lewis, Patrick; et al. "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." 2020. https://arxiv.org/abs/2005.11401
- Manning, Christopher D.; Raghavan, Prabhakar; Schütze, Hinrich. "Introduction to Information Retrieval" (Chapter 8: Evaluation in information retrieval). 2008. https://nlp.stanford.edu/IR-book/html/htmledition/evaluation-in-information-retrieval-1.html
- Zheng, Lianmin; et al. "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena." 2023. https://arxiv.org/abs/2306.05685