π Agent Memory
Descriptionβ
< What is it? >β
Agent memory is information that an application saves and retrieves to provide context for later agent steps or conversations. Short-term memory tracks the current conversation; long-term memory retains useful information across conversations.
This page explains a general memory readβwrite loop, followed by implementations using LangChain/LangGraph and LlamaIndex. The loop is framework-independent pseudocode; each framework section introduces its own APIs.
Implementationβ
< The memory readβwrite loop >β
- Load: restore the current thread's state and look up relevant memories for the authenticated user.
- Build context: combine the current request, recent history, task summary, and selected memories within the model's context budget.
- Run: let the agent answer or call tools. Record completed actions and useful results as the task progresses.
- Save: checkpoint updated state and write selected durable facts to the long-term store.
The following is pseudocode. The storage methods and agent runner represent application components, not a specific library API.
def handle_turn(user_id, thread_id, message):
# IDs come from the authenticated application session.
thread_key = (user_id, thread_id)
state = checkpoints.load(thread_key) or empty_state()
profile = memory_store.get((user_id, "profile")) or {}
notes = memory_store.search(
namespace=(user_id, "notes"), query=message, limit=3
)
context = build_context(
request=message,
history=state["messages"],
summary=state["summary"],
profile=profile,
retrieved_notes=notes,
token_budget=6000, # Illustrative budget for history and memories.
)
# Also checkpoint progress after completed tool steps inside the runner.
result = run_agent(context, state, checkpoint_key=thread_key)
checkpoints.save(thread_key, result.updated_state)
# Save selected, validated facts; update existing records when corrected.
for fact in select_durable_facts(message, result):
memory_store.upsert((user_id, "notes"), fact.key, fact.record)
return result.answer
build_context selects or summarizes older history rather than sending every saved record. Preserve complete tool-call/result pairs when trimming messages. Keep each stored fact's source and update time, support correction and deletion, and treat retrieved content as data rather than authority to override application instructions.
< Wiring memory into LangGraph >β
LangChain is a framework for building agents, and its agents run on LangGraph, which manages workflow execution and persistent state. For an existing LangGraph builder, attach a checkpointer for conversation state and a store for cross-conversation memory:
from langgraph.checkpoint.memory import InMemorySaver
from langgraph.store.memory import InMemoryStore
store = InMemoryStore()
graph = builder.compile(checkpointer=InMemorySaver(), store=store)
# An explicitly supplied preference, accessible across this user's threads.
store.put(("user-42", "profile"), "travel", {"transport": "train"})
config = {"configurable": {"thread_id": "trip-001"}}
graph.invoke(
{"messages": [{"role": "user", "content": "Help plan my trip."}]},
config,
)
This is a setup fragment: builder must already define the agent's state, nodes, and edges. Nodes must retrieve the profile from the store and pass it to the model; attaching a store does not automatically insert its contents into the prompt. Resolve user namespaces from authenticated context and authorize access to each thread.
Reuse thread_id to continue a conversation; choose a new one for a separate conversation. Both in-memory backends above lose their contents when the process ends. Use database-backed implementations, such as PostgresSaver and PostgresStore, when data must survive restarts. See LangGraph: Add memory for complete setup examples.
< Wiring memory into LlamaIndex >β
LlamaIndex is another framework for building agents and RAG applications. Its Memory class combines recent conversation history with optional long-term memory blocks. See the LlamaIndex memory guide.
| Memory | LlamaIndex component | What it keeps |
|---|---|---|
| Short-term | Memory | Recent messages within a token budget |
| Long-term: fixed information | StaticMemoryBlock | Predefined information, such as a profile loaded from your database |
| Long-term: extracted facts | FactExtractionMemoryBlock | Facts extracted from older messages, such as a travel preference |
| Long-term: searchable history | VectorMemoryBlock | Older message batches retrieved by embedding similarity |
The following setup fragment combines short-term history with extracted facts. It assumes agent and llm are configured and runs inside an async function or notebook.
from llama_index.core.memory import Memory, FactExtractionMemoryBlock
memory = Memory.from_defaults(
session_id="user-42:trip-001",
token_limit=8000,
chat_history_token_ratio=0.7,
token_flush_size=1000,
memory_blocks=[
FactExtractionMemoryBlock(
name="user_facts", llm=llm, max_facts=50
),
],
)
await agent.run("I prefer trains. Let's plan a trip to Kyoto.", memory=memory)
response = await agent.run("Which city did I mention?", memory=memory)
Reuse the memory object across turns. When recent history exceeds its allocated budget, older messages are flushed to the blocks for processing. Retrieved block content is combined with recent history for later calls. Omit memory_blocks if only short-term history is needed. See LlamaIndex's memory introduction.
- Example: βKyotoβ is available in recent history for the second call. After enough conversation triggers flushing, the fact-extraction block can retain βThe user prefers trains.β
Persistence needs explicit storage. Set async_database_uri for durable conversation storage; the default uses in-memory SQLite. Save and restore extracted facts separately, and use a persistent vector store if adding VectorMemoryBlock. For a new conversation, create a new session while reconnecting to the same user's long-term records. Keep storage and retrieval scoped to that user. See the storage configuration.
Related ideasβ
- Agentic AI System introduces agents, tools, state, and debugging.
- Agentic RAG Evaluation includes tests for memory behavior.
Referenceβ
- LangChain overview (LangChain docs)
- Memory overview (LangChain docs)
- Memory (LlamaIndex docs)
- Improved long- and short-term memory for LlamaIndex agents (LlamaIndex)
- Add memory (LangChain docs)