How it works
Documents are split into chunks, embedded and indexed ahead of time. When a question arrives, the app retrieves the most relevant chunks (with vector, keyword or hybrid search), places them in the prompt with instructions such as 'answer only from these sources and cite them', and the model writes an answer grounded in that text. The name comes from a 2020 research paper by a team at Facebook AI Research.
RAG keeps answers current without retraining: update a document and the next answer reflects it. It also allows citations and per-user permissions, since you decide what each user's search may return. Most of the quality comes from retrieval rather than the model, so chunking, search method, reranking and good source documents matter more than prompt wording. It reduces hallucination without removing it, so answers still need spot checks and a fallback when nothing relevant is found.
Retrieval-augmented generation (RAG) pros and cons
Pros
- Answers from your own, up-to-date content without training a model
- Can cite sources, so answers are easier to check
- Respects permissions when search is filtered per user
- Works with any model, so you can switch providers later
Cons
- Only as good as the retrieval; poor chunks give poor answers
- Extra moving parts: ingestion, embeddings, an index and evaluation
- Retrieved text adds input tokens to every call
- Still hallucinates sometimes, especially when sources disagree
When to use Retrieval-augmented generation (RAG)
Pick it when
- Support bots, internal knowledge search and document Q&A
- Content changes often and answers must reflect it
- Users need to see where an answer came from
Skip it when
- The model must learn a style or format rather than facts (consider fine-tuning)
- All the text it needs already fits comfortably in the prompt
Retrieval-augmented generation (RAG) vs the alternatives
Related terms
More in AI and LLMs
Search and memory