Skip to content

RAG Evaluation Explained: How to Measure Retrieval-Augmented Generation Quality

RAG evaluation splits retrieval and generation metrics, from context precision and recall to faithfulness, using frameworks like RAGAS to score pipelines.

TruLens diagram showing RAG evaluation steps with relevance and groundedness scores
TruLens evaluation workflow diagram · Credit: TruLens

RAG evaluation is the practice of measuring a retrieval-augmented generation system on two separate axes, retrieval quality and generation quality, so that a wrong or missing document and a fabricated answer built on a correct document get diagnosed as different failures. A retrieval-augmented generation pipeline works in two stages: a retriever searches a corpus for relevant documents or chunks, then a generator model writes an answer using that retrieved material as context. Scoring the finished answer alone hides which stage broke. A team chasing a low-quality output might rewrite prompts for weeks when the real defect is a retriever that never surfaced the right paragraph in the first place.

Why RAG Evaluation Splits Retrieval from Generation

Comparison matrix comparing Retrieval side and Generation side on what it checks, key metrics and failure caught

RAG evaluation treats a retrieval-augmented generation pipeline as two components that fail in different ways, so a single end-to-end score cannot tell a team whether to fix the retriever or the generator. The RAGAS paper, one of the first to formalize this problem, decomposes evaluation into retrieval metrics and generation metrics precisely because the two stages have different failure modes and different fixes (arxiv.org). A generator can produce a fluent, confident answer even when the retriever handed it the wrong context, or no useful context at all, which is exactly the scenario retrieval-augmented generation was supposed to prevent.

This split matters for a practical reason: retrieval failures and generation failures point to different remediations. Poor retrieval quality usually calls for better chunking, a different embedding model such as OpenAI's text-embedding-3 family, or a reranking step from a vendor like Cohere. A generation problem calls for prompt changes, a different base model such as Claude or GPT-4, or stricter instructions to stay within the retrieved context. Measuring retrieval quality and generation quality separately, rather than blending them into one number, turns a vague complaint about answer quality into an actionable engineering diagnosis.

Retrieval-Side Metrics: Precision, Recall, Hit Rate, MRR, and NDCG

Retrieval-side metrics answer one question: did the system fetch the evidence a correct answer would need, and how highly ranked was it? Context precision checks how much of what got retrieved was actually useful, while context recall checks how much of what was needed actually got retrieved. Those two numbers can diverge sharply, and both are read together as a single retrieval quality signal rather than in isolation. A retriever that returns twenty chunks and buries one relevant paragraph among noise scores well on context recall and poorly on context precision. A retriever that returns exactly one tightly relevant chunk but misses a second piece of evidence the question needed scores the opposite way, and either pattern signals a retrieval quality problem worth investigating before touching the generator at all.

MetricWhat It MeasuresHow It Is ComputedBest Fit
Context precisionFraction of retrieved chunks that are relevant to the queryRelevant chunks retrieved divided by total chunks retrievedCatching noisy, over-broad retrieval
Context recallFraction of the evidence needed for a correct answer that was actually retrievedEvidence retrieved divided by evidence required, against a labeled answer keyCatching missed or incomplete retrieval
Retrieval hit rateWhether at least one relevant document appears anywhere in the retrieved setA coarse pass or fail check per queryFast, low-effort sanity check
Mean reciprocal rankHow highly ranked the first relevant result isThe reciprocal of the rank position of the first relevant hit, averaged across queriesRewarding retrievers that surface the answer early
NDCGGraded-relevance ranking quality across the full result listDiscounted cumulative gain normalized against an ideal rankingCases with multiple relevant documents at varying relevance

Mean reciprocal rank and NDCG both come from classic information retrieval research and get reused for RAG because ranking still matters once retrieval feeds a generator with a limited context window. A Hugging Face cookbook walkthrough shows how to compute these retrieval metrics, including retrieval hit rate, alongside generation metrics on the same evaluation set (huggingface.co). A vector database like Pinecone or an open source store like Chroma both expose the raw similarity scores needed to compute retrieval hit rate and the ranking metrics above it. One caution applies across all five metrics: retrieving more chunks is not automatically better. Irrelevant or redundant text dilutes the context window and can make it harder for the generator to identify which retrieved passage actually answers the question, so a retrieval strategy tuned purely to maximize recall can quietly hurt the generation stage it feeds.

Generation-Side Metrics: Faithfulness, Groundedness, and Answer Relevancy

Generation-side metrics answer a different question than retrieval metrics: given the evidence the retriever supplied, did the generator produce an answer that is accurate, complete, and actually grounded in that evidence? A faithfulness score measures whether each claim in the generated answer can be traced back to the retrieved context, which makes it the primary defense against hallucination in a RAG system. Faithfulness and several related RAG metrics are commonly computed by prompting a separate language model to judge whether each claim is supported, an LLM-as-a-judge approach, rather than by a hand-built rule.

  • Faithfulness score: measures whether each claim in the generated answer traces back to the retrieved context, the core check against fabricated statements, defined alongside context precision and recall in the RAGAS paper (arxiv.org).
  • Groundedness: a closely related or often synonymous term for faithfulness, used by several evaluation frameworks to describe whether statements are supported by cited or retrieved evidence rather than pulled from the model's own parametric memory.
  • Answer relevancy: measures whether the generated answer actually addresses the question asked, independent of faithfulness, since an answer can be fully grounded in retrieved text and still fail to answer what was asked (docs.ragas.io).
  • Answer correctness: compares the generated answer against a gold reference answer when one exists, combining semantic similarity with factual overlap for a stricter correctness signal.

Citation presence alone is not proof of faithfulness. A generator can attach a plausible-looking citation to a claim the cited passage does not actually support, a risk that applies equally to a narrow internal tool and a consumer-facing assistant like Perplexity or Microsoft Copilot, so a faithfulness check needs to verify the claim against the content of the retrieved text, not just confirm that a citation exists somewhere in the answer. Treating groundedness and faithfulness as multi-dimensional rather than a single pass or fail flag catches this gap.

RAGAS and Other RAG Evaluation Frameworks

RAGAS is the most widely referenced open framework purpose-built for RAG evaluation, and several adjacent tools, including TruLens and DeepEval, implement overlapping or complementary metric sets. Choosing a framework mostly comes down to what already sits closest to the retrieval and generation stack in use, whether that is a LangChain pipeline, a LlamaIndex application, or a custom stack built directly on an API like OpenAI's or Anthropic's.

  1. RAGAS computes context precision, context recall, faithfulness, and answer relevancy from a query, answer, and context triple, largely without requiring hand-labeled gold answers for every query (arxiv.org).
  2. LlamaIndex's built-in evaluation modules provide retriever evaluators for hit rate and mean reciprocal rank plus response evaluators for faithfulness and relevancy, inside the same framework many teams use to build the pipeline itself (developers.llamaindex.ai).
  3. Hugging Face's RAG evaluation cookbook offers a practical reference implementation that combines retrieval checks and generation checks on one evaluation set (huggingface.co).
  4. Managed cloud evaluation tooling, such as the retrieval and response scoring built into Amazon Bedrock Knowledge Bases, evaluates a fully managed retriever without requiring a separate open source harness (docs.aws.amazon.com).
  5. Reranker-aware evaluation matters whenever a reranking model sits between the initial retriever and the generator, since adding or swapping a reranker changes context precision and recall and should trigger a re-measurement of the retrieval metrics above it (docs.cohere.com).

None of these frameworks compute an identical metric formula under the hood. A faithfulness score from RAGAS and a groundedness score from an Azure AI Foundry or Vertex AI evaluation pipeline can differ in how strictly they parse claims, so comparing scores across frameworks without checking the underlying computation risks a false apples-to-apples comparison.

Reference-Free vs Reference-Based RAG Evaluation

RAG evaluation methods split further into reference-free and reference-based approaches, and the choice affects both what a metric can measure and how much curation work it demands upfront. Pinecone's own evaluation guidance walks through this distinction directly (pinecone.io).

  • Reference-free evaluation scores an answer using only the query, the retrieved context, and the generated answer itself, so metrics like faithfulness and answer relevancy can run continuously on live production traffic without a pre-written gold answer.
  • Reference-based evaluation compares the generated answer or the retrieved context against a curated gold answer or gold evidence set, giving a more precise correctness signal at the cost of upfront annotation work, and only covering the specific queries included in that fixed set.
  • A practical evaluation program typically leans on reference-based scoring offline, against a curated RAG evaluation set, before shipping a pipeline change, then switches to reference-free metrics in production monitoring to catch drift as the retrieval corpus, embeddings, or user queries change over time.

Neither approach alone is sufficient for a production system. Reference-free evaluation scales to every live query but cannot catch a case where the generator is faithful to context that was subtly wrong to retrieve in the first place. Reference-based evaluation catches that case but only for the queries someone thought to annotate in advance, so most mature RAG programs run both rather than choosing one.

Building a RAG Evaluation Set

A RAG evaluation set is the fixed collection of queries, expected evidence, and often gold answers that turns RAG evaluation from a one-off spot check into a repeatable measurement. Teams building internal knowledge assistants on top of a vector database, a document store like Elasticsearch, or a managed platform like Databricks tend to learn this the hard way: building an evaluation set well matters more than picking a metric, since even the best metric produces meaningless numbers on a poorly constructed test set.

  • Representative user queries pulled from real usage patterns, not only the easy queries a pipeline is already known to handle well.
  • Known relevant documents or chunk identifiers attached to each query, so context precision and context recall can be scored against an actual answer key rather than estimated after the fact.
  • Gold answers or evidence annotations where feasible, enabling reference-based answer correctness checks alongside reference-free faithfulness scoring.
  • Adversarial or edge-case queries, including questions the corpus genuinely cannot answer, to check whether the pipeline appropriately declines rather than fabricates a confident-sounding response.
  • Periodic refresh of the evaluation set as the retrieval corpus and embeddings change, since a static test set can go stale in exactly the same way as the corpus it was built to test.

The most defensible RAG evaluation reports both component metrics, retrieval and generation scored separately, and end-to-end task success, rather than collapsing everything into one aggregate number that hides which stage of the pipeline needs attention.

Further reading

Frequently Asked Questions

What is the difference between RAG evaluation and general AI model evaluation?

RAG evaluation specifically measures a retrieval-augmented generation pipeline on two separate axes: whether the retriever fetched the right evidence, and whether the generator used that evidence correctly. General AI model evaluation covers a broader set of benchmarks and eval techniques for model capability across many tasks, not the retrieval-plus-generation architecture RAG evaluation is built around.

Can a RAG system score well on faithfulness but still give a wrong answer?

Yes. A faithfulness score only checks whether the generated claims are supported by the retrieved context, not whether that context was the right context to retrieve in the first place. A RAG pipeline can produce an answer that is fully grounded in retrieved documents and still be wrong or unhelpful if the retriever surfaced the wrong documents, which is why retrieval metrics and generation metrics are evaluated separately.

Do I need a hand-labeled gold answer set to run RAG evaluation?

Not always. Reference-free metrics like faithfulness and answer relevancy can score an answer using only the query, retrieved context, and generated response, which works well for ongoing production monitoring. Reference-based metrics, which compare against a curated gold answer or gold evidence set, give a more precise correctness signal and are worth the annotation effort for a fixed offline evaluation set used before shipping pipeline changes.

Is retrieving more context chunks always better for RAG evaluation scores?

No. Retrieving more chunks can lower generation-side scores even when retrieval metrics look fine, because irrelevant or redundant text dilutes the context window and makes it harder for the generator to identify the actually relevant evidence. A RAG evaluation set should test whether a pipeline's chunk count and ranking strategy help or hurt the generator, not assume that more retrieved evidence is automatically an improvement.

Share this guide

Julian Beaumont

Julian Beaumont covers artificial intelligence and large language models for techshooked, following the path from research paper to deployed feature. His standard is anti-hype: ask what a model actually does, what trained it, how it fails, and whether a benchmark measures what the announcement claims.