Knowledge & retrieval

Retrieval-Augmented Generation

RAG

A technique in which relevant documents are retrieved at question time and placed in the model's context, so that it answers from your information and not only from its training.

IN ADVANCE · indexingDocumentsPDFs, wiki, emailChunkinto sectionsEmbedchunk → vectorVector databaseready to searchAT QUESTION TIMEQuestionfrom the userRetrievemost relevant passagesAugmentadd them to the promptGenerateanswer with citationsThe model is not retrained; knowledge is placed in the context when the question is asked.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/retrieval-augmented-generation

In plain terms

An open-book exam. A model on its own answers from memory, and its memory stops at its training date and has never included your company's files. With RAG, the system first looks up the passages that bear on the question, hands them to the model, and asks it to answer from those pages.

Why it matters

RAG is the standard way to make a general model useful on private and current knowledge: policies, contracts, product documentation, support history. It needs no model training, updates the moment a document changes, and lets answers point to their sources. Its quality ceiling is the retrieval step: if the right passage is not found, the best model cannot save the answer.

Example

An insurer loads 12,000 pages of policy wordings into a RAG system. A call-centre agent asks whether water damage from a burst pipe is covered under a particular home policy. The system retrieves the three relevant clauses and the model answers in two sentences, with the clause numbers, in under five seconds.

Most often confused with

RAG vs. Fine-tuning

RAGSupplies knowledge at question time
Fine-tuningChanges the model's behaviour through training

RAG gives the model facts to read; fine-tuning changes how the model behaves. For knowledge that is large, private or changes often, RAG is almost always the right first choice. Fine-tuning suits teaching a style, a format or a narrow skill. The two can be combined.

Origin: Introduced by Patrick Lewis and colleagues at Facebook AI Research in 2020.

Under the hood

Two pipelines. Indexing, done ahead of time: load documents, split them into chunks, embed each chunk, store vectors and text in an index. Querying: embed the question, retrieve the top candidates (often with hybrid search), rerank, assemble a prompt with the passages and instructions to cite and to abstain when the answer is absent, then generate. Common failure points: poor chunking, retrieval that misses the passage, too many irrelevant passages, and answers that drift from the sources. Evaluate retrieval (recall, precision) and generation (faithfulness, relevance) separately. Variants include agentic RAG, GraphRAG and, for small corpora, loading everything into a long context window.

Written by Mehmet Erkek · Last updated: