π LlamaIndex
Descriptionβ
< What is it? >β
LlamaIndex (Large Language Model Index) is an open-source framework for building context-augmented LLM applications. It helps an application ingest, parse, organize, retrieve, and pass private or domain-specific data to an LLM. Common uses include RAG, document question answering, chatbots, and agentic workflows.
It is a frameworkβnot an LLM or a vector database. It provides common interfaces that connect data sources, embedding models, vector stores, LLMs, and application logic.
Key pointsβ
< RAG stages >β
Most RAG applications are commonly summarized as five broad stages: Loading, Indexing, Storing, Querying, and Evaluation. This page shows Parsing separately between Loading and Indexing because document RAG often must extract structure before chunking and embedding.
| Stage | What happens |
|---|---|
| Loading | Bring data from files, PDFs, websites, databases, or APIs into the application. LlamaIndex readers and connectors handle many source types. |
| Parsing | Extract useful structure from raw documents: text, headings, tables, images, layout, and metadata. Scanned documents may also need OCR. |
| Indexing | Chunk parsed content into nodes sized for embedding and retrieval, so the selected context can fit within the LLM's context window alongside the prompt and response. Then create embeddings, indexes, and metadata. |
| Storing | Persist indexes, vectors, documents, and metadata so the application does not need to re-index unchanged data. |
| Querying | Retrieve relevant context, optionally use sub-queries, routing, reranking, or multi-step strategies, then synthesize an LLM response. |
| Evaluation | Measure retrieval and response qualityβsuch as relevance, faithfulness, latency, and costβwhen comparing or changing a pipeline. |
Chunking is not only a storage step: one raw document may be too long to embed or send directly to an LLM. The application retrieves only the most relevant nodes and must leave context-window space for instructions, the user question, and generated output.
< Persisting and reloading an index >β
index.storage_context.persist() saves the index's configured storage data so the application can reuse it without rebuilding the index and embeddings. With LlamaIndex's default local stores, the default directory is ./storage; an explicit directory is usually clearer.
# Save the index after it has been built
index.storage_context.persist(
persist_dir="./storage/my_index"
)
# Reconstruct it later
from llama_index.core import StorageContext, load_index_from_storage
storage_context = StorageContext.from_defaults(
persist_dir="./storage/my_index"
)
index = load_index_from_storage(storage_context)
The storage context contains the document/node store, index metadata, andβin the default local vector storeβthe vectors. When using an external vector database, vectors may already persist remotely, so the exact components written depend on the configured backend. See LlamaIndex storage persistence.
< Important RAG concepts >β
| Stage | Concept | Purpose |
|---|---|---|
| Loading | Documents and nodes | A Document wraps a source, such as a PDF, API result, or database record. A Node is the atomic chunk used by LlamaIndex; it carries text or other content plus metadata and links to its source. |
| Readers / data connectors | Ingest data from files, APIs, databases, cloud drives, and other sources into Documents and Nodes. | |
| Indexing | Embeddings | Numerical representations created by an embedding model. The query is embedded too, so a vector store can find semantically similar nodes. |
| Indexing / storing | Indexes and vector stores | An index organizes data for retrieval. A vector store persists and searches embedding vectors, often with metadata filters. |
| Querying | Retrievers | Define how relevant context is selected from an index for a query. Retrieval strategy strongly affects relevance, latency, and cost. |
| Routers | Choose one or more candidate retrievers or data sources based on the query and their metadata. | |
| Node postprocessors | Transform, filter, or rerank retrieved nodes before they are sent to the LLM. | |
| Response synthesizers | Use the user query and selected nodes to generate the final LLM response. | |
| Query / chat engines | Package retrieval, prompting, and response synthesis for question answering or multi-turn conversation. |
< Agent workflow building blocks >β
These are LlamaIndex-specific building blocks for agent workflows. AgentWorkflow is a high-level class; the other items are event types or configuration used while it runs.
| Type | Item | Role |
|---|---|---|
| Class | AgentWorkflow | Coordinates one or more tool-using agents and can hand a task to another agent. A query engine or retriever can be one of its tools in agentic RAG. |
| Agent configuration | can_handoff_to | An allowlist of agent names that the current agent may hand a task to. For example, ["WriteAgent"] permits a handoff only to WriteAgent; it restricts options but does not choose the next agent. |
| Event | AgentInput | Prepared messages and the current agent for the next step. It may include the user question, history, or tool results. |
| Event | AgentOutput | The completed result of one agent step: its response and any selected tool calls. A tool call continues the workflow instead of ending it. |
| Event | AgentStream | An incremental output delta, used to render a response as it arrives rather than waiting for the step to finish. |
| Event | ToolCall | A request to run a named tool with specific arguments. It is emitted before the tool executes. |
| Event | ToolCallResult | The tool's returned output after execution, including an error when applicable. It can become context for the next agent step. |
A typical cycle is AgentInput β agent step β AgentOutput β ToolCall β ToolCallResult β AgentInput. When an AgentOutput has no tool call, it finishes the task.
< Workflow events and steps >β
LlamaIndex workflows are event-driven. A step is code that does work; an event is the typed data or signal that moves between steps. An event does not run itselfβits type routes it to one or more compatible steps.
| Aspect | Step | Event |
|---|---|---|
| What it is | A Python method registered with @step | A typed data or signal object |
| Role | Receives an event, performs work, then returns or sends another event | Carries data and determines which compatible step runs next |
| Examples | my_step, retrieve, generate_answer | StartEvent, ResultEvent, StopEvent |
workflow.run(...)
β
StartEvent βββ [my_step] βββ ResultEvent βββ [next_step] βββ StopEvent βββ _done
event step event step event internal
| Name | Meaning |
|---|---|
StartEvent | A built-in event that starts a workflow. Arguments passed to workflow.run(...) become fields on it. |
my_step | An arbitrary name for a developer-written method. The @step decorator registers it as workflow logic; it could instead be named retrieve or generate_answer. |
StopEvent | A built-in terminal event. Returning StopEvent(result=...) ends the workflow and makes its result available from await workflow.run(...). |
_done | A private LlamaIndex finalizer that may appear in a workflow diagram. It handles completion after StopEvent; application code does not define or call it. The leading underscore indicates an internal implementation detail. |
-
Multi-step loop example.
LoopEventis not built in: it is a developer-definedEventcarrying the data needed to return tostep_one. In this example,step_tworandomly loops back withLoopEventor continues withSecondEvent;step_threefinishes withStopEvent.# Example: multi-step workflows with custom eventsfrom llama_index.core.workflow import (Event,StartEvent,StopEvent,Workflow,step,)import randomclass LoopEvent(Event):first_input: strclass FirstEvent(Event):first_output: strclass SecondEvent(Event):second_output: strclass MyWorkflow(Workflow):# Step one will trigger on a StartEvent or a LoopEvent@stepasync def step_one(self, ev: StartEvent | LoopEvent) -> FirstEvent:print(ev.first_input)return FirstEvent(first_output="First step complete")# Step two returns either a SecondEvent or a LoopEvent@stepasync def step_two(self, ev: FirstEvent) -> SecondEvent | LoopEvent:print(ev.first_output)if random.randint(0, 1) == 0:print("Bad thing happened")return LoopEvent(first_input="Back to step one.")else:print("Good thing happened")return SecondEvent(second_output="Second step complete.")@stepasync def step_three(self, ev: SecondEvent) -> StopEvent:print(ev.second_output)return StopEvent(result="Workflow complete.")w = MyWorkflow(timeout=10, verbose=False)result = await w.run(first_input="Start the workflow.")print(result)
For example, a final step can return StopEvent(result="Finished"); awaiting the workflow then returns "Finished".
< Calling an LLM from a step >β
Create or inject an LLM client once, then call its asynchronous method inside a step. Events should carry serializable inputs and outputsβnot the LLM client itself.
from llama_index.core.workflow import StartEvent, StopEvent, Workflow, step
from llama_index.llms.openai import OpenAI
class AnswerWorkflow(Workflow):
llm = OpenAI(model="gpt-4.1")
@step
async def answer(self, ev: StartEvent) -> StopEvent:
prompt = f"Answer briefly: {ev.question}"
response = await self.llm.acomplete(prompt)
return StopEvent(result=str(response))
result = await AnswerWorkflow().run(question="What is RAG?")
The LLM call is await self.llm.acomplete(prompt). The question keyword argument becomes a field on StartEvent, and the response becomes the workflow result. For a production workflow, inject clients, indexes, and configuration as resources rather than putting them in events or workflow state. See the official LLM-in-a-step workflow example.
Termination signal. To define a terminal step, annotate it as async def my_step(...) -> StopEvent and return StopEvent(result=...). The annotation declares the terminal route; the returned object is the actual runtime signal. The answer step above is terminal for this reason. There is no automatic general βdoneβ condition: application logic decides whether to continue or end normally.
return NextEvent(...)continues or loops to the step that acceptsNextEvent.return StopEvent(result=...)ends the workflow normally. A customStopEventsubclass also ends it, while giving the final result a typed schema.- Typical conditions include finishing all tasks, receiving enough retrieval evidence, using the allowed number of retries, or producing a final agent answer.
- Timeout, cancellation, or a permanently failed step end the workflow abnormally; they are not normal application-level
StopEventdecisions.
See branches and loops, custom start and stop events, and workflow termination.
< What is agentic RAG? >β
Agentic RAG makes retrieval one capability inside an LLM agent loop. Instead of always using one predetermined retrieval path, the agent can decide whether to retrieve, which source or retriever to use, whether to ask a follow-up question, rerank evidence, call another tool, and when it has enough evidence to answer.
User question
β agent chooses the next action
β retrieve from a RAG tool, call another tool, or answer
β inspect the result
β optionally take another action
β grounded final answer
The word agentic does not mean the system is automatically reliable or fully autonomous. Its tool choices are model decisions, so production systems still need clear tool descriptions, access control, iteration limits, citations, logging, and evaluation.
< Build a minimal agentic RAG with LlamaIndex >β
LlamaIndex can expose a query engine as a QueryEngineTool, then give that tool to a FunctionAgent. The agent decides when to call the tool; the query engine performs ordinary RAG when it is called.
# pip install llama-index-core llama-index-llms-openai
from llama_index.core import SimpleDirectoryReader, VectorStoreIndex
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.core.tools import QueryEngineTool
from llama_index.llms.openai import OpenAI
# Build the usual RAG index.
documents = SimpleDirectoryReader("./company_docs").load_data()
index = VectorStoreIndex.from_documents(documents)
query_engine = index.as_query_engine(similarity_top_k=4)
# Give retrieval a precise name and description for the agent.
company_docs = QueryEngineTool.from_defaults(
query_engine,
name="company_docs",
description=(
"Search company policies and product documentation. "
"Use this for factual questions about the company."
),
)
agent = FunctionAgent(
llm=OpenAI(model="gpt-4o-mini"),
tools=[company_docs],
system_prompt=(
"You are a company-support agent. Always use company_docs "
"before answering a factual question about company policy. "
"If the retrieved evidence is insufficient, say so."
),
)
# In a notebook; wrap this in asyncio.run(...) in a regular Python script.
response = await agent.run(
user_msg="Can a customer cancel after the free trial ends?"
)
print(response)
To make this more agentic, give the agent several well-described tools: for example, one query engine for policies, one for product documentation, and a ticket-status or calculator tool. The agent can then route the question and make follow-up tool calls when needed.
- Use
FunctionAgentfor a compact single-agent tool loop. - Use
AgentWorkflowwhen specialist agents need to hand work to one another. - Use a custom
Workflowwhen you need deterministic branches, retry budgets, approval steps, or strict rules around retrieval.
< Minimal doc Q&A example >β
After configuring an LLM and embedding model, a small document-Q&A prototype can look like this:
from llama_index.core import SimpleDirectoryReader, VectorStoreIndex
documents = SimpleDirectoryReader("data").load_data()
index = VectorStoreIndex.from_documents(documents)
query_engine = index.as_query_engine()
response = query_engine.query("What is the refund policy?")
print(response)
This high-level API is useful for a prototype. Production systems normally customize parsing, chunking, metadata, retrieval, reranking, storage, observability, and evaluation.
< Indexing documents into Milvus >β
In a basic LlamaIndex + Milvus RAG pipeline, the path is Document β Node (chunk) β dense embedding β Milvus collection. LlamaIndex creates the dense embeddings; Milvus stores and searches them.
# pip install llama-index llama-index-vector-stores-milvus \
# llama-index-embeddings-openai
from llama_index.core import (
SimpleDirectoryReader,
StorageContext,
VectorStoreIndex,
)
from llama_index.core.node_parser import SentenceSplitter
from llama_index.embeddings.openai import OpenAIEmbedding
from llama_index.vector_stores.milvus import MilvusVectorStore
# 1. Load source files into Document objects
documents = SimpleDirectoryReader("./data").load_data()
# 2. Split Documents into smaller Node objects
splitter = SentenceSplitter(chunk_size=512, chunk_overlap=50)
nodes = splitter.get_nodes_from_documents(documents)
# 3. Configure an embedding model
embed_model = OpenAIEmbedding(model="text-embedding-3-small")
# 4. Connect LlamaIndex's vector-store adapter to Milvus
vector_store = MilvusVectorStore(
uri="http://localhost:19530",
collection_name="knowledge_base",
dim=1536, # Must match the embedding model's output dimension
)
storage_context = StorageContext.from_defaults(
vector_store=vector_store
)
# 5. Generate node embeddings and insert them into Milvus
index = VectorStoreIndex(
nodes,
storage_context=storage_context,
embed_model=embed_model,
)
Each Milvus record typically contains the embedding vector, chunk text, metadata such as its source file or page, and identifiers that connect it to the source document. To answer a question, LlamaIndex embeds the query with the same embedding model, retrieves the nearest vectors from Milvus, and supplies the retrieved chunk text to the LLM. Milvus can also add sparse/BM25 fields for hybrid search, but basic dense RAG uses LlamaIndex to generate the embeddings.
See Milvus's LlamaIndex integration guide.
< What it does not solve automatically >β
LlamaIndex does not make an answer correct merely because it retrieves documents. You still need representative source data, appropriate chunking and retrieval, access control, evaluation, and a way to show or verify sources when accuracy matters.
Comparisonβ
< Fixed RAG vs agentic RAG >β
| Aspect | Fixed RAG | Agentic RAG |
|---|---|---|
| Control flow | Application code follows a predetermined retrieve-then-answer path | An agent chooses actions during a tool-use loop |
| Retrieval | Usually one known retriever or query engine is called once | The agent may skip, repeat, reroute, or refine retrieval |
| Tools | Often retrieval only | Retrieval can be combined with search, databases, calculators, APIs, or specialist agents |
| Best for | Straightforward document Q&A | Multi-source, multi-step, or action-oriented questions |
| Cost and latency | Lower and more predictable | Higher and more variable because tool calls can repeat |
| Reliability work | Easier to constrain and evaluate | Needs stronger guardrails, limits, traces, and tool/result evaluation |
Fixed RAG can still include reranking, metadata filters, or multiple retrieval stagesβthe key distinction is whether the control flow is predefined by the application or selected dynamically by an agent.
< RAG vs LlamaIndex >β
| Item | What it is | Relationship |
|---|---|---|
| Retrieval-augmented generation (RAG) | A design pattern that retrieves external context before an LLM responds | Can be implemented with many libraries or custom code |
| LlamaIndex | A software framework for context-augmented LLM applications | Provides components to implement RAG and broader agentic workflows |
< LlamaIndex vs LangChain, LangGraph, and LangSmith >β
| Tool | Primary focus | Use it when | Relationship to LlamaIndex |
|---|---|---|---|
| LlamaIndex | Ingesting, structuring, retrieving, and using data as LLM context | Your main problem is RAG, document understanding, or data-backed agents | Provides the data and retrieval layer |
| LangChain | A configurable agent harness with model, tool, prompt, and middleware integrations | You need to compose an LLM tool loop or agent application | Overlaps with LlamaIndex; either can implement RAG or agents, and they can be combined |
| LangGraph | Low-level orchestration and runtime for long-running, stateful agent workflows | You need explicit state, branches, retries, persistence, or human approval | A LlamaIndex query engine or retriever can be one tool or node in a larger graph |
| LangSmith | Tracing, debugging, evaluation, and production monitoring | You need to inspect executions and measure whether an LLM application is improving | Complements LlamaIndex rather than replacing it; instrument the pipeline to evaluate its retrieval and responses |
They are not mutually exclusive. For example, a LangGraph workflow can call a LlamaIndex RAG tool, while LangSmith records traces and evaluates the resulting answers.
Implementationβ
- https://github.com/yujie-hao/datacamp-building-agentic-workflows-llamaindex
- Looping & branching sample project
- Concurrency & event collection
- Multi-agent system
- Workflow Self-Refection
- LLMs are capable of self-reflection: they can read their own work, critique it, and provide feedback, allowing them to take a second try when they fall short.
< Ingest JSON records into RAG >β
-
Approach: convert each JSON record into a LlamaIndex
Document. Put the information users will search for in its text, and keep identifiers, categories, and source locations in metadata.JSON recordsβDocuments: readable text + metadataβSplit long documents into chunksβGenerate embeddings β vector indexβRetrieve relevant chunks β LLM generates an answerA document keeps its text connected to its source. You can exclude selected metadata fields from embedding input while retaining them for filtering and source display. See Defining and Customizing Documents.
-
Input data: save these two policy records as
policies.json:[{"id": "returns","title": "Return Policy","category": "shopping","content": "Unused products may be returned within 30 days of delivery.","source_url": "/policies/returns"},{"id": "shipping","title": "Shipping Policy","category": "shopping","content": "Standard shipping takes 3 to 5 business days.","source_url": "/policies/shipping"}]Replace the example source paths with the locations of your actual policies.
-
Load, embed, and retrieve: install the dependencies, then run the Python code from the directory containing
policies.json. The embedding model runs locally and downloads on first use; no LLM API key is needed for this retrieval example. See LlamaIndex's Hugging Face integration.python -m pip install llama-index-core llama-index-embeddings-huggingfaceimport jsonfrom llama_index.core import Document, VectorStoreIndexfrom llama_index.core.node_parser import SentenceSplitterfrom llama_index.embeddings.huggingface import HuggingFaceEmbedding# 1. Parse JSON.with open("policies.json", encoding="utf-8") as file:records = json.load(file)# 2. Create one document per record.documents = [Document(id_=record["id"],text=f"{record['title']}\n{record['content']}",metadata={"record_id": record["id"],"category": record["category"],"source_url": record["source_url"],},# Keep identifiers and URLs out of semantic embedding input.excluded_embed_metadata_keys=["record_id", "source_url"],)for record in records]# 3. Configure an English embedding model.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5")# 4. Split longer records and build the index.index = VectorStoreIndex.from_documents(documents,embed_model=embed_model,transformations=[SentenceSplitter(chunk_size=256, chunk_overlap=30)],)# 5. Save the index locally.index.storage_context.persist(persist_dir="./storage")# 6. Retrieve evidence for a question.retriever = index.as_retriever(similarity_top_k=1)results = retriever.retrieve("How long do I have to return an unused product?")for result in results:print("Record:", result.node.metadata["record_id"])print("Text:", result.node.get_text())print("Source:", result.node.metadata["source_url"]) -
Sample retrieval output: the result includes the matching record, its text, and its source location. This output is illustrative; the retrieval example has not been executed here.
Record: returnsText: Return PolicyUnused products may be returned within 30 days of delivery.Source: /policies/returns -
Generate an answer: pass the retrieved evidence to an LLM through a query engine. The following snippet continues the example and assumes you have already initialized
llmusing a LlamaIndex LLM integration:# Assumes `llm` is an initialized LlamaIndex LLM integration.query_engine = index.as_query_engine(llm=llm,similarity_top_k=2,)response = query_engine.query("How long do I have to return an unused product?")print(response)Illustrative answer:
You can return an unused product within 30 days of delivery. -
Field mapping: choose how each field will support retrieval, filtering, or source attribution:
JSON field Treatment Title, description, policy text Include in searchable text Record ID, source URL Preserve as metadata for tracing and citations Category, language, date Preserve as metadata for filtering; optionally include in text Price, stock, other exact values Keep structured values for exact filtering or database queries -
Nested JSON: preserve field relationships and units when converting a record into readable text. For example:
{"product": {"name": "Headphones", "battery": {"hours": 30}}}Product: HeadphonesBattery life: 30 hoursYou can also embed serialized JSON. Selecting meaningful fields gives you more control over what retrieval matches.
-
Exact counts and filters: for questions such as βHow many products cost less than $100?β, use a database query or structured tool. Retrieving a few similar chunks cannot reliably calculate totals across the dataset.
-
Multilingual queries: for Chinese questions over these English documents, use a suitable multilingual embedding model for both indexing and querying, or translate the question into English before retrieval.
Related ideasβ
- Retrieval Augmented Generation (RAG) explains the retrieval pattern itself.
- Embeddings explains how text becomes vectors for semantic retrieval.
- Agentic AI System explains systems that plan and use tools over multiple steps.
Video Tutorialβ
- LangChain vs LlamaIndex
- LlamaIndex Crash Course: RAG, Nodes, Persistent Storage
Referenceβ
-
- LlamaIndex Framework documentation
- LlamaIndex origins: the GPT Index rebrand
- Introduction to RAG
- Storage persistence
- Multi-agent patterns in LlamaIndex
- Agentic RAG with LlamaIndex
- Building an agent
- Using existing tools and QueryEngineTool
- Workflows
- Calling an LLM from a workflow step
- Branches and loops
- Custom start and stop events
- Streaming output and events
- Concurrent_execution
-
LlamaIndex: Adding Personal Data to LLMs (Abid Awan | datacamp)
-
Building Agentic Workflows with LlamaIndex (datacamp)
-
What is llamaindex (geeksforgeeks)
-
LlamaIndex - series courses (Data Science Basics | Youtube)
-
Llama-Index - series courses (Prompt Engineering | Youtube)
-
Building Agentic RAG System using LlamaIndex (geeksforgeeks)
-
LlamaIndex β Did you try End to End document workflows (Shilpa Thota)
-
Retrieval-Augmented Generation (RAG) with Milvus and LlamaIndex