In plain terms
Before a model can answer from your documents, something has to pick which few paragraphs out of millions it gets to read. That picking is retrieval. It works like a research assistant who runs to the archive and comes back with the five pages most likely to hold the answer.
Why it matters
In systems built on company knowledge, retrieval decides the outcome more than the model does. A strong model handed the wrong passages gives a confident wrong answer; a modest model handed the right ones does well. When a RAG system disappoints, look at what was retrieved before blaming the model.
Example
A user asks an internal assistant, “What is our travel policy for flights over six hours?” Retrieval searches 80,000 chunks and returns five: three from the travel policy, one from an expenses FAQ, one from an outdated 2019 memo. The model's answer will be only as good as that selection, including the stale memo.
Most often confused with
Retrieval vs. Generation
Retrieval and generation are the two halves of RAG. Retrieval is a search problem with measurable right and wrong results. Generation is a writing problem. They fail in different ways and are improved with different tools, so they should be tested separately.
Under the hood
Methods: lexical search (BM25) on exact terms; dense retrieval on embeddings; hybrid combinations; and structured lookups such as SQL or graph queries. A typical pipeline retrieves a wide candidate set quickly, then narrows it with a reranker. Supporting techniques: query rewriting and expansion, metadata filters (date, department, access rights), and multi-step retrieval driven by an agent. Metrics: recall@k (was the right passage among the top k?), precision, and mean reciprocal rank. Access control belongs here: retrieval must return only what the asking user is permitted to see.