In plain terms
An open-book exam. A model on its own answers from memory, and its memory stops at its training date and has never included your company's files. With RAG, the system first looks up the passages that bear on the question, hands them to the model, and asks it to answer from those pages.
Why it matters
RAG is the standard way to make a general model useful on private and current knowledge: policies, contracts, product documentation, support history. It needs no model training, updates the moment a document changes, and lets answers point to their sources. Its quality ceiling is the retrieval step: if the right passage is not found, the best model cannot save the answer.
Example
An insurer loads 12,000 pages of policy wordings into a RAG system. A call-centre agent asks whether water damage from a burst pipe is covered under a particular home policy. The system retrieves the three relevant clauses and the model answers in two sentences, with the clause numbers, in under five seconds.
Most often confused with
RAG vs. Fine-tuning
RAG gives the model facts to read; fine-tuning changes how the model behaves. For knowledge that is large, private or changes often, RAG is almost always the right first choice. Fine-tuning suits teaching a style, a format or a narrow skill. The two can be combined.
Origin: Introduced by Patrick Lewis and colleagues at Facebook AI Research in 2020.
Under the hood
Two pipelines. Indexing, done ahead of time: load documents, split them into chunks, embed each chunk, store vectors and text in an index. Querying: embed the question, retrieve the top candidates (often with hybrid search), rerank, assemble a prompt with the passages and instructions to cite and to abstain when the answer is absent, then generate. Common failure points: poor chunking, retrieval that misses the passage, too many irrelevant passages, and answers that drift from the sources. Evaluate retrieval (recall, precision) and generation (faithfulness, relevance) separately. Variants include agentic RAG, GraphRAG and, for small corpora, loading everything into a long context window.