Skip to main content

πŸ“ Agentic RAG Evaluation

Description​

< What is it? >​

An Agentic Retrieval Augmented Generation (RAG) evaluation system runs test cases through an agent and scores both its answers and the steps used to obtain them. This approach applies to agents built with LangChain/LangGraph, LlamaIndex, or another framework.

Test cases β†’ Run agent and capture trace β†’ Score results β†’ Compare versions

The answer-quality checks also apply to a standard RAG pipeline. Agentic RAG adds checks for tool use, memory, and the sequence of steps taken.

Key points​

< Create a test dataset >​

Each case contains a question, expected answer or behavior, supporting evidence, and any required tool-use constraints. Keep reference answers hidden from the agent.

  • Example: for a fictional product, test whether the agent retrieves both plan descriptions and compares their project limits:

    case = {
    "question": "How does the Pro plan differ from Basic?",
    "reference_answer": "Basic allows 5 projects; Pro allows 50.",
    "required_evidence": ["basic-plan", "pro-plan"],
    "constraints": {
    "must_cite_sources": True,
    "max_tool_calls": 5,
    },
    }

Include ordinary questions, questions requiring multiple retrievals, missing information, tool failures, and conversations that depend on memory. The LangSmith RAG evaluation guide demonstrates the dataset–run–score workflow.

< Capture each execution >​

Record the answer, retrieved passages, tool names and arguments, tool results, memory used, latency, and token usage. Reset memory between independent cases; preserve it within an intentional multi-turn scenario.

< Score several dimensions >​

  • Evaluate components separately: a low score should point to the part that needs investigation.

    DimensionEvaluation questionPossible measurement
    Retrieval qualityDid it retrieve and rank the necessary evidence?Context precision/recall; Recall@K, NDCG, and MRR with relevance labels
    Answer correctnessDoes the answer agree with trusted reference facts?Reference-based grading
    Faithfulness / groundednessAre its claims supported by retrieved passages?Claim-support evaluation
    Answer relevanceDoes it address the user's question?Relevance grading
    CompletenessDoes it include the important information?Coverage of required facts
    CitationsDo citations support their claims, and are important claims cited?Source-ID validation, support, and coverage checks
    Multilingual behaviorIs the answer in the requested language, with facts preserved?Language checks and bilingual semantic grading
    Agent behaviorWere tool choices and arguments appropriate?Tool-call checks and task success
    Memory behaviorDid it correctly reuse or update relevant information?Multi-turn scenario checks
    EfficiencyDid it finish within the application's budget?Latency, tokens, calls, and retries
  • Faithfulness and correctness are different: an answer can accurately repeat an outdated source and still be factually wrong. A grounded answer can also omit important information.

  • Combine evaluation methods: use deterministic checks for schemas, source IDs, required actions, and budgets; use an LLM judge for semantic qualities such as correctness and support. Check its judgments against human-reviewed examples. Automated scores are useful signals, not guaranteed ground truth.

  • Match metrics to available data: some metrics need reference answers or labeled evidence; others use the question, generated answer, and retrieved context. A valid citation ID establishes that a source exists, not that it supports the claim.

Comparison​

< Tools for RAG evaluation >​

  • Choose tools for the workflow you need:

    ToolWhat it providesWhen to choose it
    RagasMetrics for faithfulness, answer relevance, retrieval, tool calls, and agent goalsCalculate quality scores in Python
    DeepEvalAutomated evaluation tests for RAG and other LLM applicationsAdd quality checks to testing or CI
    LangSmithEvaluation datasets, experiment comparisons, human review, and tracingManage evaluation and inspect failures in a platform
    Arize PhoenixOpen-source tracing and evaluation for LLM applicationsConnect quality scores with retrieval and generation traces
  • A practical starting setup: use Ragas for quality metrics and add LangSmith or Phoenix when you need experiment tracking and debugging. DeepEval is an alternative when test suites and CI checks are the priority. These tools overlap in capabilities; you do not need all four.

Implementation​

< Evaluate English sources and Chinese answers >​

  • Example evaluation record: the following contains a Chinese question, English evidence, a generated answer, and a trusted reference answer. It illustrates an application-level record, not a specific library's input schema. Keep reference answers hidden from the answering model.

    {
    "question": "ζœͺδ½Ώη”¨ηš„ε•†ε“ε―δ»₯εœ¨ε€šε°‘ε€©ε†…ι€€θ΄§οΌŸ",
    "retrieved_contexts": [
    "Unused products may be returned within 30 days of delivery."
    ],
    "generated_answer": "ζœͺδ½Ώη”¨ηš„ε•†ε“ε―δ»₯εœ¨ι€θΎΎεŽ 30 倩内退货。[S1]",
    "reference_answer": "ζœͺδ½Ώη”¨ηš„ε•†ε“ε―εœ¨ι€θΎΎεŽ30倩内退货。",
    "expected_language": "zh-CN",
    "citations": {
    "S1": "Unused products may be returned within 30 days of delivery."
    }
    }
  • Expected assessment: use a bilingual judge to check the Chinese answer against the English evidence. The table below describes the intended rubric outcome, not output captured from an evaluator run.

    CheckExpected result for this example
    Correctness and completenessPreserves 30 days, after delivery, and unused products
    FaithfulnessThese claims are supported by the English passage
    RelevanceDirectly answers the return-window question
    LanguageAnswers in Simplified Chinese
    CitationS1 exists and supports the return-policy claim

    An answer saying β€œ30 days after purchase” should fail the factual check. A language detector alone cannot verify that the meaning was preserved.

< Implement the evaluation runner >​

The following is framework-independent pseudocode; the runner and grading functions are application components.

for case in test_cases:
result = run_agent(
question=case["question"],
memory=fresh_test_memory(),
capture_trace=True,
)

scores = {
"correctness": grade_answer(
result.answer, case["reference_answer"]
),
"groundedness": grade_support(
result.answer, result.retrieved_passages
),
"retrieval": check_evidence(
result.retrieved_passages, case["required_evidence"]
),
"behavior": check_constraints(
result.answer, result.trace, case["constraints"]
),
}

save_result(case, result, scores, version="candidate")

This loop evaluates independent questions. A multi-turn case runs its sequence of messages with shared memory. Allow alternative valid tool sequences unless a particular order is required; see LangSmith's agent evaluation guidance.

< Compare and monitor >​

Run baseline and candidate versions against the same dataset and knowledge-base snapshot. Repeat cases to measure variability, inspect failures by category, and enforce explicit quality and cost thresholds. Add production failures to the regression suite and keep a separate held-out test set for final comparisons.

Use the tools in the comparison above to store traces, attach evaluation scores, and compare runs. See LangSmith evaluation types for offline testing and production monitoring.

Reference​