π Agentic RAG Evaluation
Descriptionβ
< What is it? >β
An Agentic Retrieval Augmented Generation (RAG) evaluation system runs test cases through an agent and scores both its answers and the steps used to obtain them. This approach applies to agents built with LangChain/LangGraph, LlamaIndex, or another framework.
Test cases β Run agent and capture trace β Score results β Compare versions
The answer-quality checks also apply to a standard RAG pipeline. Agentic RAG adds checks for tool use, memory, and the sequence of steps taken.
Key pointsβ
< Create a test dataset >β
Each case contains a question, expected answer or behavior, supporting evidence, and any required tool-use constraints. Keep reference answers hidden from the agent.
-
Example: for a fictional product, test whether the agent retrieves both plan descriptions and compares their project limits:
case = {"question": "How does the Pro plan differ from Basic?","reference_answer": "Basic allows 5 projects; Pro allows 50.","required_evidence": ["basic-plan", "pro-plan"],"constraints": {"must_cite_sources": True,"max_tool_calls": 5,},}
Include ordinary questions, questions requiring multiple retrievals, missing information, tool failures, and conversations that depend on memory. The LangSmith RAG evaluation guide demonstrates the datasetβrunβscore workflow.
< Capture each execution >β
Record the answer, retrieved passages, tool names and arguments, tool results, memory used, latency, and token usage. Reset memory between independent cases; preserve it within an intentional multi-turn scenario.
< Score several dimensions >β
-
Evaluate components separately: a low score should point to the part that needs investigation.
Dimension Evaluation question Possible measurement Retrieval quality Did it retrieve and rank the necessary evidence? Context precision/recall; Recall@K, NDCG, and MRR with relevance labels Answer correctness Does the answer agree with trusted reference facts? Reference-based grading Faithfulness / groundedness Are its claims supported by retrieved passages? Claim-support evaluation Answer relevance Does it address the user's question? Relevance grading Completeness Does it include the important information? Coverage of required facts Citations Do citations support their claims, and are important claims cited? Source-ID validation, support, and coverage checks Multilingual behavior Is the answer in the requested language, with facts preserved? Language checks and bilingual semantic grading Agent behavior Were tool choices and arguments appropriate? Tool-call checks and task success Memory behavior Did it correctly reuse or update relevant information? Multi-turn scenario checks Efficiency Did it finish within the application's budget? Latency, tokens, calls, and retries -
Faithfulness and correctness are different: an answer can accurately repeat an outdated source and still be factually wrong. A grounded answer can also omit important information.
-
Combine evaluation methods: use deterministic checks for schemas, source IDs, required actions, and budgets; use an LLM judge for semantic qualities such as correctness and support. Check its judgments against human-reviewed examples. Automated scores are useful signals, not guaranteed ground truth.
-
Match metrics to available data: some metrics need reference answers or labeled evidence; others use the question, generated answer, and retrieved context. A valid citation ID establishes that a source exists, not that it supports the claim.
Comparisonβ
< Tools for RAG evaluation >β
-
Choose tools for the workflow you need:
Tool What it provides When to choose it Ragas Metrics for faithfulness, answer relevance, retrieval, tool calls, and agent goals Calculate quality scores in Python DeepEval Automated evaluation tests for RAG and other LLM applications Add quality checks to testing or CI LangSmith Evaluation datasets, experiment comparisons, human review, and tracing Manage evaluation and inspect failures in a platform Arize Phoenix Open-source tracing and evaluation for LLM applications Connect quality scores with retrieval and generation traces -
A practical starting setup: use Ragas for quality metrics and add LangSmith or Phoenix when you need experiment tracking and debugging. DeepEval is an alternative when test suites and CI checks are the priority. These tools overlap in capabilities; you do not need all four.
Implementationβ
< Evaluate English sources and Chinese answers >β
-
Example evaluation record: the following contains a Chinese question, English evidence, a generated answer, and a trusted reference answer. It illustrates an application-level record, not a specific library's input schema. Keep reference answers hidden from the answering model.
{"question": "ζͺδ½Ώη¨ηεεε―δ»₯ε¨ε€ε°ε€©ε ιθ΄§οΌ","retrieved_contexts": ["Unused products may be returned within 30 days of delivery."],"generated_answer": "ζͺδ½Ώη¨ηεεε―δ»₯ε¨ιθΎΎε 30 倩ε ιθ΄§γ[S1]","reference_answer": "ζͺδ½Ώη¨ηεεε―ε¨ιθΎΎε30倩ε ιθ΄§γ","expected_language": "zh-CN","citations": {"S1": "Unused products may be returned within 30 days of delivery."}} -
Expected assessment: use a bilingual judge to check the Chinese answer against the English evidence. The table below describes the intended rubric outcome, not output captured from an evaluator run.
Check Expected result for this example Correctness and completeness Preserves 30 days, after delivery, and unused products Faithfulness These claims are supported by the English passage Relevance Directly answers the return-window question Language Answers in Simplified Chinese Citation S1exists and supports the return-policy claimAn answer saying β30 days after purchaseβ should fail the factual check. A language detector alone cannot verify that the meaning was preserved.
< Implement the evaluation runner >β
The following is framework-independent pseudocode; the runner and grading functions are application components.
for case in test_cases:
result = run_agent(
question=case["question"],
memory=fresh_test_memory(),
capture_trace=True,
)
scores = {
"correctness": grade_answer(
result.answer, case["reference_answer"]
),
"groundedness": grade_support(
result.answer, result.retrieved_passages
),
"retrieval": check_evidence(
result.retrieved_passages, case["required_evidence"]
),
"behavior": check_constraints(
result.answer, result.trace, case["constraints"]
),
}
save_result(case, result, scores, version="candidate")
This loop evaluates independent questions. A multi-turn case runs its sequence of messages with shared memory. Allow alternative valid tool sequences unless a particular order is required; see LangSmith's agent evaluation guidance.
< Compare and monitor >β
Run baseline and candidate versions against the same dataset and knowledge-base snapshot. Repeat cases to measure variability, inspect failures by category, and enforce explicit quality and cost thresholds. Add production failures to the regression suite and keep a separate held-out test set for final comparisons.
Use the tools in the comparison above to store traces, attach evaluation scores, and compare runs. See LangSmith evaluation types for offline testing and production monitoring.
Related ideasβ
-
Normalized Discounted Cumulative Gain (NDCG) evaluates the ordering of relevant evidence.
-
Mean Reciprocal Rank (MRR) measures how early the first relevant passage appears.
-
Agentic RAG Observability records the execution data that evaluators can score.
-
Agentic AI System covers memory, tool use, and debugging.
-
Retrieval Augmented Generation (RAG) explains retrieval and generation.
Referenceβ
-
Evaluate a RAG application (LangSmith docs)
-
Evaluation types (LangSmith docs)
-
Evaluating AI agents at the run, trace, and thread level (LangChain)
-
Available evaluation metrics (Ragas docs)