LlamaIndex
LlamaIndex is a framework focused on connecting LLMs with private data, the mainstream choice for RAG applications. More "focused" than LangChain with cleaner API design.
Core abstractions
- Document: raw data from any source (PDF, webpage, DB row).
- Node: minimal unit after chunking, with metadata.
- Index: organization of nodes (vector index / keyword index / knowledge graph index / tree index).
- QueryEngine: executes a query against an index + LLM.
- ResponseSynthesizer: composes retrieved nodes into the final answer.
Workflow
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
documents = SimpleDirectoryReader("data").load_data()
index = VectorStoreIndex.from_documents(documents)
query_engine = index.as_query_engine()
response = query_engine.query("What does the doc say?")
Advanced features
- Query Pipelines: chain retrieval + rerank + LLM into debuggable pipelines.
- Agents / Tools: built-in ReAct, OpenAI Function Calling agents.
- Multi-modal: native support for images and PDF tables.
- Workflows: event-driven orchestration (v0.10+).
vs LangChain
| Dimension | LlamaIndex | LangChain |
|---|---|---|
| Core focus | RAG / data access | Agent / Chain orchestration |
| API elegance | High | Medium (historical baggage) |
| Ecosystem breadth | Medium (focused) | Very wide (full-stack) |
| Learning curve | Gentle | Steep |
Newcomers get a better RAG experience starting with LlamaIndex; complex agents go with LangChain / LangGraph.