You evaluate a RAG system by scoring four layers separately (source data, retrieval, generation, and the end-to-end system) against a fixed test set, then tracking the scores over time. RAG (retrieval-augmented generation) is a design where a retriever finds passages in a knowledge base and a large language model (LLM) answers from them.
Scoring layers separately matters because one end-to-end score hides the failing component. A long-context RAG study from Google Cloud AI Research, tested with Gemma-2-9B-Chat on Natural Questions, found that "RAG accuracy consistently falls below retrieval recall for both retrievers, indicating that the presence of relevant information does not guarantee correct answers."
The Four Layers at a Glance
Each layer answers one question about your pipeline, in the order data flows through it. The common three-part RAG triad scores retrieval and generation, while the data layer adds a check on the corpus itself.
| Layer | Question it answers | Example metrics | Needs labels? |
|---|---|---|---|
| Data | Is the right, current, clean content in the index? | Freshness, coverage, ingestion cleanliness | Coverage needs gold passages; freshness and cleanliness do not |
| Retrieval | Do the right passages rank near the top? | Recall@k, MRR, nDCG, context precision | Yes for Recall@k, MRR, and nDCG; optional for context precision |
| Generation | Is the answer supported, on-topic, and correct? | Faithfulness, answer relevance, correctness | Only for correctness |
| System | Is the product fast, affordable, and safe? | Latency, cost per query, refusal rate | Only for unanswerable questions used in refusal rate |
"Needs labels" means a person must first write reference answers or mark which documents are relevant. Metrics without labels can be computed by code or graded by an LLM judge.
Step 1: Build an Evaluation Dataset
A golden dataset is a fixed set of real questions, each paired with a reference answer and the source passages that support it. Every later step scores your pipeline against this set, so build it first.
Start with 50 to 100 questions taken from real user logs. That range is a starting point your team can grow, not an industry standard.
Build the question mix from four groups:
- Real questions: Pull them from search logs, chat transcripts, or support tickets so the wording matches how users ask.
- Synthetic questions: Generate extra questions with an LLM to widen topic coverage, then drop any that look nothing like production queries.
- Unanswerable questions: Add questions your corpus cannot answer, and label the expected response as a refusal.
- Time-sensitive questions: Add questions about prices, versions, or policies whose correct answer changes over time.
Compare synthetic questions with real logs by phrasing, length, and topic before you keep them, since generated wording can drift from how users type.
Time-sensitive reference answers need a check against current sources before you label them. One option is the Olostep Answers endpoint, which searches the live web and returns a source-grounded answer you can compare with your label. Store a review date on each of these questions.
Keep the Dataset Fresh
Reference answers expire when the source pages behind them change, so re-check and re-label the golden dataset on a schedule. Pick a cadence that matches how often your sources change.
This problem is documented for technical content. A 2026 preprint, the retrieval benchmark drift study, notes that "temporal changes in technical corpora, such as API deprecations and code reorganizations, can render existing benchmarks stale."
Re-label outside the schedule whenever a monitored source changes. Store the source URL with every golden question, so a changed page maps straight to the questions that need review.
Step 2: Evaluate the Data Layer First
The data layer is the corpus your retriever searches, meaning the pages, documents, and chunks in your index. If the answer is missing, stale, or buried in boilerplate, no retriever or model can fix it.
Check this layer first because its errors look like retrieval failures or hallucinations (claims with no support in the sources) further down. For corpora built from websites, web scraping for RAG is the step where missing pages, old copies, and page clutter can enter the index.
Coverage: Does the Answer Exist in the Index?
Coverage measures whether a supporting passage for each golden question exists anywhere in your corpus. Search the index for each question's gold passage by source URL or exact text, then report coverage as the share of questions with a match.
Treat a missing passage as a coverage failure, not a retrieval failure. No ranking change can surface text that was never indexed, so tuning the retriever here wastes time.
The usual fix is to crawl the full source site instead of a hand-picked list of URLs. The Olostep Crawl endpoint follows links across a site and returns each page as Markdown that you can chunk and index.
Freshness: Is the Content Still Current?
Freshness measures how much of your index is older than your freshness window, which is the maximum document age you accept. Track two numbers: the share of indexed documents older than that window, and how often answers change after you re-ingest the sources.
Outdated passages cause confident wrong answers. On the HoH outdated-information benchmark (ACL 2025, open LLMs tested on Wikipedia snapshots), "even when current information is successfully retrieved, the mere presence of outdated information in the context leads to at least 20% performance drop in mainstream LLMs."
Generation metrics will not catch this failure. Faithfulness checks whether the answer matches the retrieved context, so an answer that repeats an outdated passage still scores as grounded.
The fix is to re-ingest pages when they change. Olostep's Monitors API for change detection watches source pages and reports changes, which you can use to trigger re-ingestion and a new evaluation run.
Ingestion Quality: Is the Text Clean?
Ingestion quality measures whether your chunks contain the page's useful content and nothing else. Sample a few dozen chunks at random and check each one for these problems:
- Navigation menus, sidebars, and footers
- Cookie banners and pop-up text
- Tables flattened into run-on text
- Empty or near-empty pages whose content loads through JavaScript
Turn those checks into two simple ratios. Boilerplate ratio is the share of tokens in sampled chunks that come from menus, banners, and footers. Table-survival rate is the share of source tables that keep intact rows and columns after ingestion.
Re-check both ratios whenever you change your scraper or parser. Olostep's guide to LLM-ready data describes what clean source text for models looks like, which gives you a target for the sample review.
Step 3: Evaluate Retrieval
Retrieval evaluation checks whether the right passages appear near the top of the results for each golden question. Run it on the golden dataset after the data layer passes, so a low score points at the retriever.
Two families of metrics apply here. Label-based metrics compare results with known relevant documents, while LLM-judged metrics ask a model to grade the retrieved chunks.
Label-Based Metrics: Recall@k, MRR, and nDCG
Label-based metrics score ranked results against documents you have already marked as relevant. The Databricks retrieval metrics docs define Recall@k as "The fraction of known relevant documents that appear in the top-k results."
Take one golden question with two relevant documents, A and B. The retriever returns five results in this order: C, A, D, E, F.
- Recall@5: One of the two relevant documents appears in the top 5, so Recall@5 is 0.5.
- MRR (mean reciprocal rank): The first relevant result sits at rank 2, so this question scores 1/2. MRR averages that score across all questions.
- nDCG (normalized discounted cumulative gain): Each result gets a graded label, such as 2 for "fully answers" and 1 for "partly answers." Lower ranks count less, and 1.0 is the ideal order.
Pick the metric that matches how your generator uses results. Recall@k fits answers that need several facts, MRR fits answers that depend on the first hit, and nDCG fits teams with graded labels.
LLM-Judged Metrics: Context Precision and Context Recall
Context precision and context recall use an LLM judge, a language model prompted to grade outputs with a rubric, to score retrieved chunks. Context precision asks whether the relevant chunks are present and ranked high. Context recall asks whether the chunks contain every fact the reference answer needs.
Calibrate both metrics against a small hand-labeled sample before you trust them. Skipping calibration leaves you unable to tell a retrieval drop from judge drift; the next step's judge section covers the method.
Chunking choices move both scores. Small chunks can split one fact across pieces and lower context recall, while large chunks can pull in unrelated text and lower context precision. Compare semantic chunking strategies on the same golden set and keep the one that holds up on both metrics.
Step 4: Evaluate Generation
Generation evaluation checks whether the answer is supported by the retrieved context, on-topic, and correct. Run it on questions where retrieval passed, so a low score points at the prompt or model.
The RAG triad scores three links in the chain. Add answer correctness when you have reference answers.
- Context relevance: The retrieved context relates to the question. Step 3 already scores this side.
- Faithfulness: Every claim in the answer is supported by the retrieved context. Some tools call this groundedness.
- Answer relevance: The answer addresses the question that was asked.
- Answer correctness: The answer matches the reference answer in your golden dataset. This metric sits outside the triad and needs labels.
Low faithfulness is the clearest sign of hallucination in this layer. Common fixes are instructions to answer only from the context, a citation for each claim, and search-grounded LLM responses for questions your static index does not cover.
Using an LLM as a Judge
An LLM judge scores answers with a written rubric, for example "1 if every claim is supported, 0 otherwise." Judges grade thousands of answers quickly, but their scores need checks first.
- Hand-label a sample of answers with the same rubric the judge uses.
- Fix the judge model version and prompt so scores stay comparable across runs.
- Re-run the judge on the same answers several times and measure how much scores vary.
- Report agreement with humans together with the sample size.
Published agreement figures come from small samples too. In a Dell case-aware LLM-judge study (a 2026 preprint on 60 sampled turns from one enterprise support domain), "Agreement between the LLM judge and the majority human vote was: Hallucination: 88%". That figure describes one small sample in one domain, so measure agreement on your own data.
Step 5: Evaluate the Full System
System evaluation measures what users experience end to end: speed, cost, refusals, and safety. Run these checks on the full pipeline in a staging environment.
- Latency: Track p50 (the median response time) and p95 (the time 95% of requests finish under). The p95 number shows the slow tail that users notice.
- Cost per query: Add embedding, retrieval, reranking, and LLM token costs for each request.
- Refusal rate: Measure how often the system declines unanswerable golden questions. Also count false refusals on answerable ones.
- Safety: Test prompt injection, meaning instructions hidden in user input or retrieved pages that try to override your system prompt.
Prompt injection matters more when the corpus comes from the open web, since any scraped page can carry hidden instructions. Add a few test pages with planted instructions to a staging index and check that the model ignores them.
Step 6: Diagnose Failures and Automate the Loop
Diagnose a failed question by checking the layers in order: data, retrieval, generation, then system. The first layer that fails is usually the root cause.
| Symptom | Likely layer | Check | Fix |
|---|---|---|---|
| Answer cites an old price or policy | Data (freshness) | Compare the indexed copy with the live page | Re-ingest the page and monitor the source |
| "I don't know" on an answerable question | Data (coverage) | Search the index for the gold passage | Crawl the missing pages |
| Chunks contain menus or cookie text | Data (ingestion) | Review boilerplate ratio on sampled chunks | Extract main content as clean Markdown |
| Gold passage is indexed but not in the top k | Retrieval | Recall@k for that question | Tune chunking, embeddings, or reranking |
| Right chunks retrieved, answer adds claims | Generation | Faithfulness score | Restrict the prompt to context and require citations |
| Grounded answer misses the question | Generation | Answer relevance score | Rewrite the prompt or query handling |
| Confident answer to an unanswerable question | System | Refusal rate | Add refusal instructions or a retrieval score cutoff |
| Slow or costly responses | System | p95 latency and cost per query | Lower k, cache results, or use a smaller model |
Run the evaluation in CI (continuous integration) before every deploy, and block the deploy when a score falls below its threshold. Set starter thresholds from your first baseline run, then raise them as the pipeline improves.
The example below shows a scorecard run in the style of Ragas. The threshold values are placeholders for your team to set.
# illustrative; check your framework's current API
from ragas import evaluate
from ragas.metrics import faithfulness, answer_relevancy, context_precision, context_recall
# Each row: question, pipeline answer, retrieved contexts, reference answer
dataset = load_golden_dataset("golden.jsonl") # your own loader
scores = evaluate(
dataset,
metrics=[faithfulness, answer_relevancy, context_precision, context_recall],
)
# Starter thresholds chosen by your team, not industry standards
thresholds = {"faithfulness": 0.85, "answer_relevancy": 0.80,
"context_precision": 0.70, "context_recall": 0.75}
failed = {m: scores[m] for m, t in thresholds.items() if scores[m] < t}
if failed:
raise SystemExit(f"Eval gate failed: {failed}")Re-run the same scorecard when sources change, as well as when code changes. A change alert triggers re-ingestion, re-ingestion triggers the eval run, and failed golden questions go to a person for re-labeling.
Choosing an Evaluation Tool
Choose an evaluation tool by where you want scores to run: in notebooks, in CI tests, or on production traces. The six tools below differ mainly in where they run and how much they host for you.
| Tool | Type | Good fit when |
|---|---|---|
| Ragas | Open-source Python library | You want ready-made RAG metrics and synthetic test set generation |
| DeepEval | Open-source Python library | You want pytest-style LLM tests that run in CI |
| TruLens | Open-source library | You want feedback functions for RAG checks with tracing |
| Arize Phoenix | Open-source observability tool | You want to trace requests and run evals on those traces |
| LangSmith | Platform from LangChain (cloud or self-hosted) | You build with LangChain or LangGraph and want datasets, tracing, and experiments in one place |
| Braintrust | Hosted platform | You want experiment tracking, datasets, and production logging with scoring |
Olostep is not an evaluation tool or an LLM judge. It is the web data layer that crawls sources, returns clean Markdown, and monitors pages for changes. That work keeps the corpus these tools score complete and current.
FAQ
How are RAG pipelines evaluated?
RAG pipelines are evaluated by scoring each layer (source data, retrieval, generation, and the full system) against a golden dataset of real questions. Teams re-run those scores before each deploy and after source content changes.
What are the key metrics used to evaluate RAG performance?
Data metrics are coverage, freshness, and ingestion cleanliness, while retrieval metrics are Recall@k, MRR, nDCG, context precision, and context recall. Generation uses faithfulness, answer relevance, and correctness, and the system layer uses latency, cost per query, and refusal rate.
What is the RAG triad?
The RAG triad is a set of three checks: context relevance, faithfulness (groundedness), and answer relevance. It covers retrieval and generation but does not test whether the corpus is complete or current.
How do I evaluate RAG without ground-truth answers?
Use reference-free metrics such as faithfulness, answer relevance, and context precision, graded by an LLM judge. Then hand-label a small sample to confirm the judge agrees with people.
How many test questions do I need?
Start with 50 to 100 real questions from user logs, a starting point your team chooses rather than a standard. Add questions each time you find a new type of failure.
Can LLM-judged metrics replace labeled relevance?
LLM-judged metrics can cover far more questions than hand labels, but they need a labeled sample for calibration. Label-based metrics such as Recall@k also return the same score on every run, which makes them a steadier check for retrieval.
How do I know if my knowledge base is out of date?
Measure the share of indexed documents older than your freshness window, and re-ingest a sample to see how many answers change. Monitoring source pages for changes flags stale documents when the source updates.



