In plain terms
Classic RAG finds the passages that look most like the question. That works for “what is the notice period in this contract?” and fails for “what are the main themes across these 5,000 complaints?”, because no single passage holds that answer. GraphRAG reads everything in advance, notes who and what is mentioned and how they connect, groups related items and writes a summary of each group. Broad questions are then answered from those summaries.
Why it matters
It matters when your questions are about the whole picture or about chains of connections, as in investigations, audits, research reviews and market intelligence. It is also expensive: building the graph means running a language model over every chunk of every document, and the graph must be kept current. For questions whose answer sits in one or two passages, classic RAG is cheaper and just as good. Start there and add the graph only when broad questions keep failing.
Example
An audit team loads 3,000 internal audit reports. Asked “Which control weaknesses recur across branches?”, classic RAG returns five passages from five reports and a summary of those five. GraphRAG returns six recurring themes, the number of reports in which each appears, and links to the reports, because it had already clustered and summarised the whole collection.
Most often confused with
GraphRAG vs. Classic RAG
Both give the model outside knowledge at question time. Classic RAG treats the collection as a pile of separate passages; GraphRAG first works out how the contents relate. The first is the default for looking up specific facts. The second earns its extra cost on multi-step questions and on questions about the collection as a whole.
Origin: The term comes from Microsoft Research, which published the approach and released it as open source in 2024.
Under the hood
In the approach published by Microsoft Research, indexing has four stages: split documents into chunks; have a language model extract entities, relationships and claims from each chunk; merge them into a graph; run community detection (the Leiden algorithm) to form a hierarchy of clusters and write a summary for each. Querying has two main modes: global search, which answers broad questions by map-reduce over community summaries, and local search, which starts from the entities in the question and gathers their neighbourhood and source text. Later variants reduce cost, for example LazyGraphRAG, which defers summarisation to query time. The name is also used loosely for any RAG over a knowledge graph. Watch indexing cost, incremental updates and the quality of entity resolution.