Skip to main content

πŸ“ Agentic RAG Observability

Description​

< What is it? >​

Observability is the ability to understand a system's behavior from recorded execution data. In an Agentic Retrieval Augmented Generation (RAG) application, it helps developers inspect model calls, retrievals, tool calls, memory access, errors, and performance.

An observability system collects this data and makes it available through searchable traces, dashboards, and alerts. It can support applications built with LangChain/LangGraph, LlamaIndex, or other frameworks. See LangSmith's observability concepts.

  • Example: an agent gives an incorrect plan comparison. Observability shows that it retrieved only the Basic plan and then generated an answer. Evaluation checks the answer against the expected comparison and marks it incomplete.

Key points​

< Traces, logs, and metrics >​

RecordWhat it capturesExample
TraceRelated operations for one request, linked by a trace IDRead memory β†’ choose a tool β†’ retrieve documents β†’ generate an answer
SpanOne operation within a trace, with timing and statusA retrieval call and its returned document IDs
Log or eventA specific occurrence, ideally linked to a trace or spanA tool timeout or a retry
MetricAn aggregate measurement over many operationsError rate, token usage, or 95th-percentile request latency

A conversation ID connects requests across turns. Preserve parent–child relationships when agents call tools or delegate work so a failed operation can be traced back to its request. LangSmith calls individual traced operations runs and groups conversation traces into threads. See its data model.

< What should a RAG system record? >​

StageUseful fields
RequestTrace ID, conversation ID, application version, question, start/end time
MemoryRecord IDs read or updated, relevant state changes, context included
RetrievalSearch query, filters, index version, document/chunk IDs, scores, selected passages
Model callModel and prompt versions, actual input messages, output, available token usage
Tool callTool name, arguments, result or error, duration, retry count
Final responseAnswer, citations, completion status, total latency, estimated cost

Capture the passages actually supplied to the model as well as the retrieved candidates. This makes it possible to distinguish retrieval failures from information lost during filtering or context construction. Store only the content needed for diagnosis, redact credentials and sensitive fields, and define retention and access rules.

Comparison​

< Observability vs evaluation >​

AspectObservabilityEvaluation
Main questionWhat happened during execution?How well did the system perform?
InputRuntime events, model/tool inputs and outputs, timingsAnswers, evidence, traces, criteria, and optional reference answers
OutputTraces, logs, dashboards, alertsScores, pass/fail results, quality comparisons
ExampleThe agent retrieved three documents and took four secondsThe answer omitted a required fact
When usedDevelopment and productionOffline tests and online production scoring

The two systems work together. Evaluators can score recorded traces, and an observability dashboard can display those scores alongside latency and errors. A successful tool call or a low error rate alone does not establish answer quality. See LangSmith evaluation types.

Implementation​

< Build the collection pipeline >​

Instrument the application β†’ Collect traces and metrics β†’ Store and search β†’ Inspect dashboards and alerts

  1. Create a trace for each request and attach a conversation ID and application version.
  2. Instrument memory, retrieval, model, and tool operations as child spans. Record errors and retries as well as successful results.
  3. Export records to an observability backend. Use framework integrations for supported calls and manual instrumentation for custom code.
  4. Monitor trends such as latency, empty retrievals, failures, repeated tool calls, and token usage. Set thresholds based on the application's requirements.
  5. Send selected traces to evaluators and turn confirmed failures into regression cases.

The following is framework-independent pseudocode for recording a retrieval tool. The span API represents an instrumentation layer, not a specific SDK.

def traced_search(query, filters, trace):
with trace.span("retrieval") as span:
span.record_input(query=query, filters=filters)
try:
documents = retriever.search(query, filters=filters)
except Exception as error:
span.record_error(type(error).__name__)
raise

span.record_output(
document_ids=[doc.id for doc in documents],
result_count=len(documents),
)
return documents

The span context manager records duration and completion status. Apply equivalent instrumentation to other operations, and propagate trace context into asynchronous or delegated work. Recording trace data does not itself grade the returned documents.

< LangChain and LlamaIndex integrations >​

LangChain is a framework for building agents; LangGraph provides workflow execution and state management. LangSmith is a separate observability and evaluation platform with tracing integrations for these frameworks. The parent page includes a LangSmith tracing setup.

LlamaIndex is a framework for agents and RAG applications. Its instrumentation captures events and spans that supported integrations can send to an observability backend. Choose an integration and configure its exporter using the LlamaIndex observability guide.

Troubleshoot​

< Common observability problems >​

SymptomCheck
Only the final answer appearsWhether retrieval, model, and custom tool calls are instrumented
Child operations appear as unrelated tracesTrace-context propagation across async tasks, processes, and services
The latest traces are missingExporter errors, sampling configuration, and flushing before process exit
Everything reports success but answers are wrongAdd evaluation of correctness, evidence support, and task completion

Reference​