In plain terms
Classic RAG works in one shot: it searches once and answers from whatever comes back. Agentic RAG works like a researcher. It splits the question into parts, chooses which source to consult for each, reads what it finds, notices what is missing or contradictory, and goes back with a better query. It answers when it has enough evidence.
Why it matters
Real questions are rarely answered by one search. Comparisons, questions that span several systems and questions whose first results are misleading all need more than one attempt. Agentic RAG raises answer quality on these. The price is more model calls, longer waits and less predictable behaviour, so it needs step limits and good logging. Simple lookups do not need it.
Example
An analyst asks, “How did churn in Germany change after the price increase, and what did customers say?” The agent queries the data warehouse for churn figures, searches support tickets for price complaints, sees that the tickets it found pre-date the increase, narrows the date range and searches again, then writes an answer citing both sources. Four searches and forty seconds, where classic RAG would have made one search in four.
Most often confused with
Agentic RAG vs. Classic RAG
In classic RAG the retrieval steps are written by developers and always run the same way. In agentic RAG retrieval is a tool in the model's hands. The first is faster, cheaper and easier to test. The second copes with questions the developers did not foresee.
Under the hood
Retrieval is exposed to the model as tools: vector or hybrid search over document sets, SQL over databases, graph queries, web search. Common behaviours: query decomposition, routing each sub-question to the right source, grading retrieved passages for relevance, rewriting queries, and checking the draft answer against the evidence. Named patterns include Self-RAG and corrective RAG. Controls: a cap on steps and spend, stopping rules, and compaction of intermediate results so the context stays clean. Evaluate the path (which searches, in what order) as well as the final answer. Since retrieved text enters the agent's context, it is a route for indirect prompt injection.