RAG (Retrieval-Augmented Generation)
Retrieval-Augmented Generation (RAG) is a paradigm that connects external knowledge bases to LLMs. The model doesn't rely solely on parametric memory — it first "retrieves" relevant chunks from the knowledge base, then "generates" the answer.
Why RAG is needed
- Knowledge freshness: model weights are frozen post-training, RAG dynamically supplements.
- Traceability: cite specific sources, reduces hallucination.
- Domain customization: use enterprise private data without retraining the model.
Core flow
- Offline indexing: chunk documents → embedding → store in vector DB.
- Online retrieval: user question → embedding → top-k similar chunks from vector DB.
- Generation: feed retrieved chunks as context to LLM, let it answer based on those.
Advanced topics
- Hybrid retrieval: vector retrieval + BM25 keyword retrieval fusion.
- Rerank: stronger cross-encoder to refine initial ranking.
- Query rewriting: let LLM rewrite user questions for better retrieval.
- Multi-hop retrieval: a question needs to retrieve multiple referencing chunks.