π Agentic RAG Observability
Descriptionβ
< What is it? >β
Observability is the ability to understand a system's behavior from recorded execution data. In an Agentic Retrieval Augmented Generation (RAG) application, it helps developers inspect model calls, retrievals, tool calls, memory access, errors, and performance.
An observability system collects this data and makes it available through searchable traces, dashboards, and alerts. It can support applications built with LangChain/LangGraph, LlamaIndex, or other frameworks. See LangSmith's observability concepts.
- Example: an agent gives an incorrect plan comparison. Observability shows that it retrieved only the Basic plan and then generated an answer. Evaluation checks the answer against the expected comparison and marks it incomplete.
Key pointsβ
< Traces, logs, and metrics >β
| Record | What it captures | Example |
|---|---|---|
| Trace | Related operations for one request, linked by a trace ID | Read memory β choose a tool β retrieve documents β generate an answer |
| Span | One operation within a trace, with timing and status | A retrieval call and its returned document IDs |
| Log or event | A specific occurrence, ideally linked to a trace or span | A tool timeout or a retry |
| Metric | An aggregate measurement over many operations | Error rate, token usage, or 95th-percentile request latency |
A conversation ID connects requests across turns. Preserve parentβchild relationships when agents call tools or delegate work so a failed operation can be traced back to its request. LangSmith calls individual traced operations runs and groups conversation traces into threads. See its data model.
< What should a RAG system record? >β
| Stage | Useful fields |
|---|---|
| Request | Trace ID, conversation ID, application version, question, start/end time |
| Memory | Record IDs read or updated, relevant state changes, context included |
| Retrieval | Search query, filters, index version, document/chunk IDs, scores, selected passages |
| Model call | Model and prompt versions, actual input messages, output, available token usage |
| Tool call | Tool name, arguments, result or error, duration, retry count |
| Final response | Answer, citations, completion status, total latency, estimated cost |
Capture the passages actually supplied to the model as well as the retrieved candidates. This makes it possible to distinguish retrieval failures from information lost during filtering or context construction. Store only the content needed for diagnosis, redact credentials and sensitive fields, and define retention and access rules.
Comparisonβ
< Observability vs evaluation >β
| Aspect | Observability | Evaluation |
|---|---|---|
| Main question | What happened during execution? | How well did the system perform? |
| Input | Runtime events, model/tool inputs and outputs, timings | Answers, evidence, traces, criteria, and optional reference answers |
| Output | Traces, logs, dashboards, alerts | Scores, pass/fail results, quality comparisons |
| Example | The agent retrieved three documents and took four seconds | The answer omitted a required fact |
| When used | Development and production | Offline tests and online production scoring |
The two systems work together. Evaluators can score recorded traces, and an observability dashboard can display those scores alongside latency and errors. A successful tool call or a low error rate alone does not establish answer quality. See LangSmith evaluation types.
Implementationβ
< Build the collection pipeline >β
Instrument the application β Collect traces and metrics β Store and search β Inspect dashboards and alerts
- Create a trace for each request and attach a conversation ID and application version.
- Instrument memory, retrieval, model, and tool operations as child spans. Record errors and retries as well as successful results.
- Export records to an observability backend. Use framework integrations for supported calls and manual instrumentation for custom code.
- Monitor trends such as latency, empty retrievals, failures, repeated tool calls, and token usage. Set thresholds based on the application's requirements.
- Send selected traces to evaluators and turn confirmed failures into regression cases.
The following is framework-independent pseudocode for recording a retrieval tool. The span API represents an instrumentation layer, not a specific SDK.
def traced_search(query, filters, trace):
with trace.span("retrieval") as span:
span.record_input(query=query, filters=filters)
try:
documents = retriever.search(query, filters=filters)
except Exception as error:
span.record_error(type(error).__name__)
raise
span.record_output(
document_ids=[doc.id for doc in documents],
result_count=len(documents),
)
return documents
The span context manager records duration and completion status. Apply equivalent instrumentation to other operations, and propagate trace context into asynchronous or delegated work. Recording trace data does not itself grade the returned documents.
< LangChain and LlamaIndex integrations >β
LangChain is a framework for building agents; LangGraph provides workflow execution and state management. LangSmith is a separate observability and evaluation platform with tracing integrations for these frameworks. The parent page includes a LangSmith tracing setup.
LlamaIndex is a framework for agents and RAG applications. Its instrumentation captures events and spans that supported integrations can send to an observability backend. Choose an integration and configure its exporter using the LlamaIndex observability guide.
Troubleshootβ
< Common observability problems >β
| Symptom | Check |
|---|---|
| Only the final answer appears | Whether retrieval, model, and custom tool calls are instrumented |
| Child operations appear as unrelated traces | Trace-context propagation across async tasks, processes, and services |
| The latest traces are missing | Exporter errors, sampling configuration, and flushing before process exit |
| Everything reports success but answers are wrong | Add evaluation of correctness, evidence support, and task completion |
Related ideasβ
- Agentic RAG Evaluation measures answer quality, tool use, memory behavior, and efficiency.
- Agentic AI System provides a debugging workflow.
- Agent Memory explains the state and stored records an agent reads and writes.
- Retrieval Augmented Generation (RAG) explains retrieval and generation.
Referenceβ
- Observability concepts (LangSmith docs)
- Trace LangChain applications (LangSmith docs)
- Evaluation types (LangSmith docs)
- Observability (LlamaIndex docs)