Skip to main content

πŸ“ LlamaIndex

Description​

< What is it? >​

LlamaIndex (Large Language Model Index) is an open-source framework for building context-augmented LLM applications. It helps an application ingest, parse, organize, retrieve, and pass private or domain-specific data to an LLM. Common uses include RAG, document question answering, chatbots, and agentic workflows.

It is a frameworkβ€”not an LLM or a vector database. It provides common interfaces that connect data sources, embedding models, vector stores, LLMs, and application logic.

Key points​

< RAG stages >​

Most RAG applications are commonly summarized as five broad stages: Loading, Indexing, Storing, Querying, and Evaluation. This page shows Parsing separately between Loading and Indexing because document RAG often must extract structure before chunking and embedding.

StageWhat happens
LoadingBring data from files, PDFs, websites, databases, or APIs into the application. LlamaIndex readers and connectors handle many source types.
ParsingExtract useful structure from raw documents: text, headings, tables, images, layout, and metadata. Scanned documents may also need OCR.
IndexingChunk parsed content into nodes sized for embedding and retrieval, so the selected context can fit within the LLM's context window alongside the prompt and response. Then create embeddings, indexes, and metadata.
StoringPersist indexes, vectors, documents, and metadata so the application does not need to re-index unchanged data.
QueryingRetrieve relevant context, optionally use sub-queries, routing, reranking, or multi-step strategies, then synthesize an LLM response.
EvaluationMeasure retrieval and response qualityβ€”such as relevance, faithfulness, latency, and costβ€”when comparing or changing a pipeline.

Chunking is not only a storage step: one raw document may be too long to embed or send directly to an LLM. The application retrieves only the most relevant nodes and must leave context-window space for instructions, the user question, and generated output.

< Persisting and reloading an index >​

index.storage_context.persist() saves the index's configured storage data so the application can reuse it without rebuilding the index and embeddings. With LlamaIndex's default local stores, the default directory is ./storage; an explicit directory is usually clearer.

# Save the index after it has been built
index.storage_context.persist(
persist_dir="./storage/my_index"
)

# Reconstruct it later
from llama_index.core import StorageContext, load_index_from_storage

storage_context = StorageContext.from_defaults(
persist_dir="./storage/my_index"
)
index = load_index_from_storage(storage_context)

The storage context contains the document/node store, index metadata, andβ€”in the default local vector storeβ€”the vectors. When using an external vector database, vectors may already persist remotely, so the exact components written depend on the configured backend. See LlamaIndex storage persistence.

< Important RAG concepts >​

StageConceptPurpose
LoadingDocuments and nodesA Document wraps a source, such as a PDF, API result, or database record. A Node is the atomic chunk used by LlamaIndex; it carries text or other content plus metadata and links to its source.
Readers / data connectorsIngest data from files, APIs, databases, cloud drives, and other sources into Documents and Nodes.
IndexingEmbeddingsNumerical representations created by an embedding model. The query is embedded too, so a vector store can find semantically similar nodes.
Indexing / storingIndexes and vector storesAn index organizes data for retrieval. A vector store persists and searches embedding vectors, often with metadata filters.
QueryingRetrieversDefine how relevant context is selected from an index for a query. Retrieval strategy strongly affects relevance, latency, and cost.
RoutersChoose one or more candidate retrievers or data sources based on the query and their metadata.
Node postprocessorsTransform, filter, or rerank retrieved nodes before they are sent to the LLM.
Response synthesizersUse the user query and selected nodes to generate the final LLM response.
Query / chat enginesPackage retrieval, prompting, and response synthesis for question answering or multi-turn conversation.

< Agent workflow building blocks >​

These are LlamaIndex-specific building blocks for agent workflows. AgentWorkflow is a high-level class; the other items are event types or configuration used while it runs.

TypeItemRole
ClassAgentWorkflowCoordinates one or more tool-using agents and can hand a task to another agent. A query engine or retriever can be one of its tools in agentic RAG.
Agent configurationcan_handoff_toAn allowlist of agent names that the current agent may hand a task to. For example, ["WriteAgent"] permits a handoff only to WriteAgent; it restricts options but does not choose the next agent.
EventAgentInputPrepared messages and the current agent for the next step. It may include the user question, history, or tool results.
EventAgentOutputThe completed result of one agent step: its response and any selected tool calls. A tool call continues the workflow instead of ending it.
EventAgentStreamAn incremental output delta, used to render a response as it arrives rather than waiting for the step to finish.
EventToolCallA request to run a named tool with specific arguments. It is emitted before the tool executes.
EventToolCallResultThe tool's returned output after execution, including an error when applicable. It can become context for the next agent step.

A typical cycle is AgentInput β†’ agent step β†’ AgentOutput β†’ ToolCall β†’ ToolCallResult β†’ AgentInput. When an AgentOutput has no tool call, it finishes the task.

< Workflow events and steps >​

LlamaIndex workflows are event-driven. A step is code that does work; an event is the typed data or signal that moves between steps. An event does not run itselfβ€”its type routes it to one or more compatible steps.

AspectStepEvent
What it isA Python method registered with @stepA typed data or signal object
RoleReceives an event, performs work, then returns or sends another eventCarries data and determines which compatible step runs next
Examplesmy_step, retrieve, generate_answerStartEvent, ResultEvent, StopEvent
workflow.run(...)
↓
StartEvent ──→ [my_step] ──→ ResultEvent ──→ [next_step] ──→ StopEvent ──→ _done
event step event step event internal
NameMeaning
StartEventA built-in event that starts a workflow. Arguments passed to workflow.run(...) become fields on it.
my_stepAn arbitrary name for a developer-written method. The @step decorator registers it as workflow logic; it could instead be named retrieve or generate_answer.
StopEventA built-in terminal event. Returning StopEvent(result=...) ends the workflow and makes its result available from await workflow.run(...).
_doneA private LlamaIndex finalizer that may appear in a workflow diagram. It handles completion after StopEvent; application code does not define or call it. The leading underscore indicates an internal implementation detail.
  • Multi-step loop example. LoopEvent is not built in: it is a developer-defined Event carrying the data needed to return to step_one. In this example, step_two randomly loops back with LoopEvent or continues with SecondEvent; step_three finishes with StopEvent.

    # Example: multi-step workflows with custom events
    from llama_index.core.workflow import (
    Event,
    StartEvent,
    StopEvent,
    Workflow,
    step,
    )
    import random


    class LoopEvent(Event):
    first_input: str


    class FirstEvent(Event):
    first_output: str


    class SecondEvent(Event):
    second_output: str


    class MyWorkflow(Workflow):

    # Step one will trigger on a StartEvent or a LoopEvent
    @step
    async def step_one(self, ev: StartEvent | LoopEvent) -> FirstEvent:
    print(ev.first_input)
    return FirstEvent(first_output="First step complete")

    # Step two returns either a SecondEvent or a LoopEvent
    @step
    async def step_two(self, ev: FirstEvent) -> SecondEvent | LoopEvent:
    print(ev.first_output)
    if random.randint(0, 1) == 0:
    print("Bad thing happened")
    return LoopEvent(first_input="Back to step one.")
    else:
    print("Good thing happened")
    return SecondEvent(second_output="Second step complete.")

    @step
    async def step_three(self, ev: SecondEvent) -> StopEvent:
    print(ev.second_output)
    return StopEvent(result="Workflow complete.")


    w = MyWorkflow(timeout=10, verbose=False)
    result = await w.run(first_input="Start the workflow.")
    print(result)

For example, a final step can return StopEvent(result="Finished"); awaiting the workflow then returns "Finished".

< Calling an LLM from a step >​

Create or inject an LLM client once, then call its asynchronous method inside a step. Events should carry serializable inputs and outputsβ€”not the LLM client itself.

from llama_index.core.workflow import StartEvent, StopEvent, Workflow, step
from llama_index.llms.openai import OpenAI


class AnswerWorkflow(Workflow):
llm = OpenAI(model="gpt-4.1")

@step
async def answer(self, ev: StartEvent) -> StopEvent:
prompt = f"Answer briefly: {ev.question}"
response = await self.llm.acomplete(prompt)
return StopEvent(result=str(response))


result = await AnswerWorkflow().run(question="What is RAG?")

The LLM call is await self.llm.acomplete(prompt). The question keyword argument becomes a field on StartEvent, and the response becomes the workflow result. For a production workflow, inject clients, indexes, and configuration as resources rather than putting them in events or workflow state. See the official LLM-in-a-step workflow example.

Termination signal. To define a terminal step, annotate it as async def my_step(...) -> StopEvent and return StopEvent(result=...). The annotation declares the terminal route; the returned object is the actual runtime signal. The answer step above is terminal for this reason. There is no automatic general β€œdone” condition: application logic decides whether to continue or end normally.

  • return NextEvent(...) continues or loops to the step that accepts NextEvent.
  • return StopEvent(result=...) ends the workflow normally. A custom StopEvent subclass also ends it, while giving the final result a typed schema.
  • Typical conditions include finishing all tasks, receiving enough retrieval evidence, using the allowed number of retries, or producing a final agent answer.
  • Timeout, cancellation, or a permanently failed step end the workflow abnormally; they are not normal application-level StopEvent decisions.

See branches and loops, custom start and stop events, and workflow termination.

< What is agentic RAG? >​

Agentic RAG makes retrieval one capability inside an LLM agent loop. Instead of always using one predetermined retrieval path, the agent can decide whether to retrieve, which source or retriever to use, whether to ask a follow-up question, rerank evidence, call another tool, and when it has enough evidence to answer.

User question
β†’ agent chooses the next action
β†’ retrieve from a RAG tool, call another tool, or answer
β†’ inspect the result
β†’ optionally take another action
β†’ grounded final answer

The word agentic does not mean the system is automatically reliable or fully autonomous. Its tool choices are model decisions, so production systems still need clear tool descriptions, access control, iteration limits, citations, logging, and evaluation.

< Build a minimal agentic RAG with LlamaIndex >​

LlamaIndex can expose a query engine as a QueryEngineTool, then give that tool to a FunctionAgent. The agent decides when to call the tool; the query engine performs ordinary RAG when it is called.

# pip install llama-index-core llama-index-llms-openai
from llama_index.core import SimpleDirectoryReader, VectorStoreIndex
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.core.tools import QueryEngineTool
from llama_index.llms.openai import OpenAI

# Build the usual RAG index.
documents = SimpleDirectoryReader("./company_docs").load_data()
index = VectorStoreIndex.from_documents(documents)
query_engine = index.as_query_engine(similarity_top_k=4)

# Give retrieval a precise name and description for the agent.
company_docs = QueryEngineTool.from_defaults(
query_engine,
name="company_docs",
description=(
"Search company policies and product documentation. "
"Use this for factual questions about the company."
),
)

agent = FunctionAgent(
llm=OpenAI(model="gpt-4o-mini"),
tools=[company_docs],
system_prompt=(
"You are a company-support agent. Always use company_docs "
"before answering a factual question about company policy. "
"If the retrieved evidence is insufficient, say so."
),
)

# In a notebook; wrap this in asyncio.run(...) in a regular Python script.
response = await agent.run(
user_msg="Can a customer cancel after the free trial ends?"
)
print(response)

To make this more agentic, give the agent several well-described tools: for example, one query engine for policies, one for product documentation, and a ticket-status or calculator tool. The agent can then route the question and make follow-up tool calls when needed.

  • Use FunctionAgent for a compact single-agent tool loop.
  • Use AgentWorkflow when specialist agents need to hand work to one another.
  • Use a custom Workflow when you need deterministic branches, retry budgets, approval steps, or strict rules around retrieval.

< Minimal doc Q&A example >​

After configuring an LLM and embedding model, a small document-Q&A prototype can look like this:

from llama_index.core import SimpleDirectoryReader, VectorStoreIndex

documents = SimpleDirectoryReader("data").load_data()
index = VectorStoreIndex.from_documents(documents)
query_engine = index.as_query_engine()

response = query_engine.query("What is the refund policy?")
print(response)

This high-level API is useful for a prototype. Production systems normally customize parsing, chunking, metadata, retrieval, reranking, storage, observability, and evaluation.

< Indexing documents into Milvus >​

In a basic LlamaIndex + Milvus RAG pipeline, the path is Document β†’ Node (chunk) β†’ dense embedding β†’ Milvus collection. LlamaIndex creates the dense embeddings; Milvus stores and searches them.

# pip install llama-index llama-index-vector-stores-milvus \
# llama-index-embeddings-openai

from llama_index.core import (
SimpleDirectoryReader,
StorageContext,
VectorStoreIndex,
)
from llama_index.core.node_parser import SentenceSplitter
from llama_index.embeddings.openai import OpenAIEmbedding
from llama_index.vector_stores.milvus import MilvusVectorStore

# 1. Load source files into Document objects
documents = SimpleDirectoryReader("./data").load_data()

# 2. Split Documents into smaller Node objects
splitter = SentenceSplitter(chunk_size=512, chunk_overlap=50)
nodes = splitter.get_nodes_from_documents(documents)

# 3. Configure an embedding model
embed_model = OpenAIEmbedding(model="text-embedding-3-small")

# 4. Connect LlamaIndex's vector-store adapter to Milvus
vector_store = MilvusVectorStore(
uri="http://localhost:19530",
collection_name="knowledge_base",
dim=1536, # Must match the embedding model's output dimension
)
storage_context = StorageContext.from_defaults(
vector_store=vector_store
)

# 5. Generate node embeddings and insert them into Milvus
index = VectorStoreIndex(
nodes,
storage_context=storage_context,
embed_model=embed_model,
)

Each Milvus record typically contains the embedding vector, chunk text, metadata such as its source file or page, and identifiers that connect it to the source document. To answer a question, LlamaIndex embeds the query with the same embedding model, retrieves the nearest vectors from Milvus, and supplies the retrieved chunk text to the LLM. Milvus can also add sparse/BM25 fields for hybrid search, but basic dense RAG uses LlamaIndex to generate the embeddings.

See Milvus's LlamaIndex integration guide.

< What it does not solve automatically >​

LlamaIndex does not make an answer correct merely because it retrieves documents. You still need representative source data, appropriate chunking and retrieval, access control, evaluation, and a way to show or verify sources when accuracy matters.

Comparison​

< Fixed RAG vs agentic RAG >​

AspectFixed RAGAgentic RAG
Control flowApplication code follows a predetermined retrieve-then-answer pathAn agent chooses actions during a tool-use loop
RetrievalUsually one known retriever or query engine is called onceThe agent may skip, repeat, reroute, or refine retrieval
ToolsOften retrieval onlyRetrieval can be combined with search, databases, calculators, APIs, or specialist agents
Best forStraightforward document Q&AMulti-source, multi-step, or action-oriented questions
Cost and latencyLower and more predictableHigher and more variable because tool calls can repeat
Reliability workEasier to constrain and evaluateNeeds stronger guardrails, limits, traces, and tool/result evaluation

Fixed RAG can still include reranking, metadata filters, or multiple retrieval stagesβ€”the key distinction is whether the control flow is predefined by the application or selected dynamically by an agent.

< RAG vs LlamaIndex >​

ItemWhat it isRelationship
Retrieval-augmented generation (RAG)A design pattern that retrieves external context before an LLM respondsCan be implemented with many libraries or custom code
LlamaIndexA software framework for context-augmented LLM applicationsProvides components to implement RAG and broader agentic workflows

< LlamaIndex vs LangChain, LangGraph, and LangSmith >​

ToolPrimary focusUse it whenRelationship to LlamaIndex
LlamaIndexIngesting, structuring, retrieving, and using data as LLM contextYour main problem is RAG, document understanding, or data-backed agentsProvides the data and retrieval layer
LangChainA configurable agent harness with model, tool, prompt, and middleware integrationsYou need to compose an LLM tool loop or agent applicationOverlaps with LlamaIndex; either can implement RAG or agents, and they can be combined
LangGraphLow-level orchestration and runtime for long-running, stateful agent workflowsYou need explicit state, branches, retries, persistence, or human approvalA LlamaIndex query engine or retriever can be one tool or node in a larger graph
LangSmithTracing, debugging, evaluation, and production monitoringYou need to inspect executions and measure whether an LLM application is improvingComplements LlamaIndex rather than replacing it; instrument the pipeline to evaluate its retrieval and responses

They are not mutually exclusive. For example, a LangGraph workflow can call a LlamaIndex RAG tool, while LangSmith records traces and evaluates the resulting answers.

Implementation​

< Ingest JSON records into RAG >​

  • Approach: convert each JSON record into a LlamaIndex Document. Put the information users will search for in its text, and keep identifiers, categories, and source locations in metadata.

    JSON records
    ↓
    Documents: readable text + metadata
    ↓
    Split long documents into chunks
    ↓
    Generate embeddings β†’ vector index
    ↓
    Retrieve relevant chunks β†’ LLM generates an answer

    A document keeps its text connected to its source. You can exclude selected metadata fields from embedding input while retaining them for filtering and source display. See Defining and Customizing Documents.

  • Input data: save these two policy records as policies.json:

    [
    {
    "id": "returns",
    "title": "Return Policy",
    "category": "shopping",
    "content": "Unused products may be returned within 30 days of delivery.",
    "source_url": "/policies/returns"
    },
    {
    "id": "shipping",
    "title": "Shipping Policy",
    "category": "shopping",
    "content": "Standard shipping takes 3 to 5 business days.",
    "source_url": "/policies/shipping"
    }
    ]

    Replace the example source paths with the locations of your actual policies.

  • Load, embed, and retrieve: install the dependencies, then run the Python code from the directory containing policies.json. The embedding model runs locally and downloads on first use; no LLM API key is needed for this retrieval example. See LlamaIndex's Hugging Face integration.

    python -m pip install llama-index-core llama-index-embeddings-huggingface
    import json

    from llama_index.core import Document, VectorStoreIndex
    from llama_index.core.node_parser import SentenceSplitter
    from llama_index.embeddings.huggingface import HuggingFaceEmbedding

    # 1. Parse JSON.
    with open("policies.json", encoding="utf-8") as file:
    records = json.load(file)

    # 2. Create one document per record.
    documents = [
    Document(
    id_=record["id"],
    text=f"{record['title']}\n{record['content']}",
    metadata={
    "record_id": record["id"],
    "category": record["category"],
    "source_url": record["source_url"],
    },
    # Keep identifiers and URLs out of semantic embedding input.
    excluded_embed_metadata_keys=["record_id", "source_url"],
    )
    for record in records
    ]

    # 3. Configure an English embedding model.
    embed_model = HuggingFaceEmbedding(
    model_name="BAAI/bge-small-en-v1.5"
    )

    # 4. Split longer records and build the index.
    index = VectorStoreIndex.from_documents(
    documents,
    embed_model=embed_model,
    transformations=[
    SentenceSplitter(chunk_size=256, chunk_overlap=30)
    ],
    )

    # 5. Save the index locally.
    index.storage_context.persist(persist_dir="./storage")

    # 6. Retrieve evidence for a question.
    retriever = index.as_retriever(similarity_top_k=1)
    results = retriever.retrieve(
    "How long do I have to return an unused product?"
    )

    for result in results:
    print("Record:", result.node.metadata["record_id"])
    print("Text:", result.node.get_text())
    print("Source:", result.node.metadata["source_url"])
  • Sample retrieval output: the result includes the matching record, its text, and its source location. This output is illustrative; the retrieval example has not been executed here.

    Record: returns
    Text: Return Policy
    Unused products may be returned within 30 days of delivery.
    Source: /policies/returns
  • Generate an answer: pass the retrieved evidence to an LLM through a query engine. The following snippet continues the example and assumes you have already initialized llm using a LlamaIndex LLM integration:

    # Assumes `llm` is an initialized LlamaIndex LLM integration.
    query_engine = index.as_query_engine(
    llm=llm,
    similarity_top_k=2,
    )

    response = query_engine.query(
    "How long do I have to return an unused product?"
    )
    print(response)

    Illustrative answer:

    You can return an unused product within 30 days of delivery.
  • Field mapping: choose how each field will support retrieval, filtering, or source attribution:

    JSON fieldTreatment
    Title, description, policy textInclude in searchable text
    Record ID, source URLPreserve as metadata for tracing and citations
    Category, language, datePreserve as metadata for filtering; optionally include in text
    Price, stock, other exact valuesKeep structured values for exact filtering or database queries
  • Nested JSON: preserve field relationships and units when converting a record into readable text. For example:

    {"product": {"name": "Headphones", "battery": {"hours": 30}}}
    Product: Headphones
    Battery life: 30 hours

    You can also embed serialized JSON. Selecting meaningful fields gives you more control over what retrieval matches.

  • Exact counts and filters: for questions such as β€œHow many products cost less than $100?”, use a database query or structured tool. Retrieving a few similar chunks cannot reliably calculate totals across the dataset.

  • Multilingual queries: for Chinese questions over these English documents, use a suitable multilingual embedding model for both indexing and querying, or translate the question into English before retrieval.

Video Tutorial​

Reference​