A retrieval-augmented generation (RAG) prototype is quick to build: embed some documents, retrieve the nearest chunks, pass them to a model. It answers the three questions in the demo, and everyone agrees it works.
The problem is that "it answered the demo questions" is not a measurement. It says nothing about the questions users will actually ask, the documents that will be added next month, or what happens when someone changes the chunk size or the model. Most RAG systems that disappoint in production were never tested in a way that could have caught the problem.
This note covers:
- why a working prototype is weak evidence;
- the three things worth measuring, and what each failure looks like;
- how to turn those measurements into a test that blocks a bad change in CI;
- the limits of using an LLM as the judge.
Why a prototype is weak evidence
A demo is evaluated on a handful of hand-picked questions over clean, curated documents. Production differs on every one of those points:
- The documents are messier. Scanned PDFs, outdated wiki pages and two versions of the same policy all end up in the index. Retrieval starts returning chunks that are similar to the question without containing the answer, and the model fills the gap with something plausible.
- The index keeps changing. New documents shift what the nearest neighbours of a query are. A retriever that behaved well on the initial corpus can quietly degrade as the corpus grows.
- Longer context is not free. Adding more retrieved chunks does not reliably help: models use information at the start and end of a long context better than information in the middle (Liu et al., 2024).
- Every change is a regression risk. Chunk size, embedding model, re-ranker, prompt template, generation model: each one changes the output, and without a test nobody knows whether it changed for the better.
Security issues such as prompt injection through retrieved documents are real too, but they need their own controls. The quality metrics below will not detect them.
The fix is the one software engineering already uses for code: define what "correct" means, check it automatically, and refuse to deploy a change that makes it worse.
Three things to measure
A RAG answer can go wrong at three distinct points, and each one needs its own check. Looking only at the final answer hides where it went wrong, which is the information you need to fix it.
| Question | Typical metric | What a failure looks like | Where to look |
|---|---|---|---|
| Did retrieval find the information needed? | Context recall / context relevance | The answer is wrong or vague because the right chunk was never retrieved, or was buried under irrelevant ones | Chunking, embeddings, filters, re-ranking |
| Is every claim supported by the retrieved context? | Faithfulness (groundedness) | The answer states a number, date or rule that appears in none of the retrieved chunks | Prompt, generation model, context formatting |
| Does the answer address the question? | Answer relevance / factual correctness against a reference | A grounded but evasive, incomplete or off-topic answer | Query understanding, prompt, retrieval scope |
TruLens calls this set of checks the RAG triad (context relevance, groundedness, answer relevance). Ragas exposes similar ideas under different names (context recall and precision, faithfulness, response relevancy, factual correctness). The names differ; the three questions are the same.
Faithfulness deserves the strictest threshold. An answer that is grounded but incomplete is annoying. An answer that invents a contract term or a dosage is a liability. Ragas computes it by splitting the answer into individual claims and asking a judge model whether each claim can be inferred from the retrieved context (Es et al., 2024).
Keeping the three scores separate is what makes a regression actionable: a drop in context recall points to the indexing side, while a drop in faithfulness with stable recall points to the prompt or the model.
Start with a golden dataset
Every metric above needs a set of test questions. This golden dataset is the most important part of the setup, and the part most often neglected.
- Use real questions. Take them from support tickets, search logs or interviews with the people who will use the system, not only from the engineer who built it.
- Write reference answers with a domain expert. Reference-based metrics are only as good as the references.
- Include the hard cases on purpose: questions whose answer is spread across two documents, questions about outdated policies, and questions the corpus cannot answer. For the last group, the correct behaviour is to say so.
- Version it with the code. Store it in the repository (for example as JSON Lines) so that a change to the tests is reviewed like any other change.
- Grow it from failures. Every bad answer reported in production becomes a new test case.
A few dozen well-chosen questions are enough to start. Size matters less than coverage of the ways the system actually fails.
Turning evaluation into a CI gate
The goal is simple to state: no change to the prompt, retriever or model reaches production if it lowers the scores below an agreed threshold.
Below is a minimal version as a pytest test. It runs every golden question through the pipeline, scores the results with Ragas, and fails the build if a metric's average drops below its threshold. answer() stands in for your own pipeline, returning the response and the retrieved chunks.
# tests/test_rag_quality.py
import json
from pathlib import Path
import pytest
from langchain_openai import ChatOpenAI
from ragas import EvaluationDataset, evaluate
from ragas.llms import LangchainLLMWrapper
from ragas.metrics import FactualCorrectness, Faithfulness, LLMContextRecall
from rag_app.pipeline import answer # question -> (response, retrieved_contexts)
GOLDEN = Path("eval/golden.jsonl")
# Starting thresholds, to be tuned on your own data.
# Grounding is held to a stricter standard than the other two.
GATES = [
(Faithfulness(), 0.90), # claims supported by the retrieved context
(LLMContextRecall(), 0.80), # retrieval found what the reference needs
(FactualCorrectness(), 0.70), # response agrees with the reference answer
]
def build_dataset() -> EvaluationDataset:
rows = []
for line in GOLDEN.read_text(encoding="utf-8").splitlines():
item = json.loads(line)
response, contexts = answer(item["question"])
rows.append(
{
"user_input": item["question"],
"retrieved_contexts": contexts,
"response": response,
"reference": item["reference"],
}
)
return EvaluationDataset.from_list(rows)
@pytest.fixture(scope="session")
def scores():
# Pin the judge to a dated model snapshot, and keep it fixed.
judge = LangchainLLMWrapper(ChatOpenAI(model="gpt-4o-2024-08-06", temperature=0))
result = evaluate(
dataset=build_dataset(),
metrics=[metric for metric, _ in GATES],
llm=judge,
)
return result.to_pandas()
@pytest.mark.parametrize("metric, threshold", GATES, ids=lambda x: getattr(x, "name", str(x)))
def test_quality_gate(scores, metric, threshold):
mean = scores[metric.name].mean()
assert mean >= threshold, f"{metric.name} = {mean:.3f}, below the {threshold} gate"A few details make the difference between a gate people trust and one they learn to ignore:
Pin the Ragas version
Metric names and signatures change between releases. The code above is written for Ragas 0.3; version 0.4 still runs it with deprecation warnings, and replaces
evaluate()with the@experiment()decorator andLangchainLLMWrapperwithllm_factory()(migration guide). Pin the version in your lock file so that an upgrade is a deliberate, reviewed change.Set thresholds from a baseline, not from intuition
Run the suite on the current production version first, then set each threshold slightly below its score. The gate should catch regressions, not fail on the first day.
Look at the failing rows, not just the average
An average can hide a handful of very bad answers. Print the lowest-scoring questions in the test report; they are usually more informative than the mean.
Keep the evaluation affordable
Every run calls the judge model several times per question. Run the full suite on merge requests that touch the pipeline, and a smaller smoke set on everything else.
Where the tools fit
The tools overlap, but they are strongest at different stages.
| Ragas | TruLens | Prompt flow | |
|---|---|---|---|
| What it is | Python library of RAG and LLM evaluation metrics | Instrumentation and feedback functions for LLM apps, with a dashboard | Microsoft's open-source tool for building and evaluating LLM flows |
| Strongest at | Offline evaluation on a dataset, CI assertions | Tracing and scoring live or recorded traffic | Teams already building on Azure's AI platform |
| Use it for | The regression gate above | Watching quality after deployment | Orchestrating flows and their evaluations in one place |
In practice, a common split is offline evaluation in CI (a library like Ragas against the golden dataset) plus online monitoring (TruLens or an equivalent tracing tool, scoring a sample of real traffic). Offline tests tell you whether a change is safe to ship. Online monitoring tells you whether the world has changed underneath a system you did not touch.
The limits of LLM-as-judge
All the metrics above rely on a model grading another model. That is what makes them cheap enough to run on every change, and it is also their main weakness.
- Judges have biases. Studies of LLM judges report position bias, a preference for longer answers and a preference for their own outputs (Zheng et al., 2023). Use a judge from a different model family than your generator when you can.
- A judge change breaks comparability. Scores from two different judge models are not on the same scale. Pin the judge to a dated snapshot and treat a judge upgrade like a migration: re-baseline before comparing.
- Scores are noisy. Two runs on the same data rarely give identical numbers. Leave a margin between the baseline and the threshold, or the gate will fail at random.
- Humans stay in the loop. Have a person review a small random sample of judged answers on a regular schedule, especially the ones the judge scored highly. This is how you find out where the judge is wrong.
Automated scores are a regression signal, not a certificate of correctness.
Evaluation also makes cost decisions possible
Teams often stay on the largest available model because switching feels risky and nobody can prove a cheaper one is good enough. With a golden dataset and per-metric scores, that becomes a measurable question:
- Run the same suite against the current model and a cheaper candidate (a smaller hosted model or a self-hosted open-weight one).
- Compare metric by metric. A candidate may match faithfulness and fall behind only on multi-document questions: a specific gap you can work on with better retrieval or prompting, not a reason to reject it outright.
- Weigh the quality difference against measured cost and latency on your own traffic.
- Ship the switch through the same gate as any other change.
The price and speed gap between models varies too much by provider and workload to quote a general figure. The point is that you can measure it instead of guessing.
Checklist
- A versioned golden dataset of real questions, with reference answers and unanswerable cases.
- Separate scores for retrieval, grounding and answer quality.
- A CI test that blocks changes below baseline-derived thresholds, with the worst rows in the report.
- A pinned evaluation library and a pinned judge model.
- Sampled human review of the judge's decisions.
- Online monitoring on real traffic after deployment.
References
- Es, S., James, J., Espinosa-Anke, L., & Schockaert, S. (2024). RAGAS: Automated Evaluation of Retrieval Augmented Generation. EACL 2024, System Demonstrations.
- Chen, J., Lin, H., Han, X., & Sun, L. (2024). Benchmarking Large Language Models in Retrieval-Augmented Generation. AAAI 2024.
- Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the Middle: How Language Models Use Long Contexts. Transactions of the ACL.
- Zheng, L., Chiang, W.-L., Sheng, Y., et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023, Datasets and Benchmarks.
- Ragas documentation. Metric definitions and the evaluation API.
- TruLens documentation. Feedback functions and the RAG triad.
- Prompt flow. Microsoft's open-source repository and documentation.