RAG & Knowledge Systems · 6 min
GraphRAG: Retrieval Over Entities and Relationships
Flat chunk retrieval breaks on multi-hop and whole-corpus questions; GraphRAG indexes an entity graph and community summaries so query time becomes a walk, not a lucky guess.
You have already built the standard pipeline: chunk the docs, embed each chunk, store the vectors, and at query time pull back the top-k nearest neighbors and stuff them into the prompt. It works, and for a large class of questions it is all you need. This lesson is about the questions where it quietly fails, and the machinery you would add to fix them. That machinery is expensive and often unnecessary, so the real skill is knowing when it earns its keep.
Where flat retrieval runs out
Vector search answers local questions: the ones whose answer sits in a passage or two that land close to the query in embedding space. "What was the deductible on the 2022 policy?" pulls back the clause that states it. Fine.
Now try three other shapes:
- Multi-hop. "Which of our vendors share a board member with a company we've flagged for sanctions risk?" The answer lives in no single chunk. It is a join across facts scattered through different documents, and the linking entity (the board member) may never co-occur with the query terms.
- "Summarize everything about X." X is spread across 400 chunks. Top-k with k=10 hands you ten of them, and probably the ten phrased most like your query, which means ten that say roughly the same thing. Coverage and similarity pull against each other.
- Global / thematic. "What are the main themes across this whole corpus?" There is no nearest neighbor to a theme. The answer is a property of the entire dataset, not of any passage inside it. Embedding similarity has nothing to grab onto.
The failure is structural. Flat retrieval treats the corpus as a bag of independent passages ranked by similarity to the query. But the fact you need often lives in the relationships between passages, or in the aggregate, and neither survives chunking. Cranking k up to 100 does not rescue you: you blow the context window, bury the signal, and still miss facts that sit nowhere near the query lexically.
What GraphRAG does instead
The Microsoft paper (Edge et al., 2024) reframes the hard case as query-focused summarization over a corpus, then builds an index that captures structure. Two phases.
Index time (expensive, done once). Walk every chunk through an LLM and extract entities (people, orgs, systems, concepts) plus the relationships between them, each with a short description. Merge duplicate mentions across documents into single nodes. The result is a knowledge graph: nodes are entities, edges are relationships, and both carry LLM-written descriptions with provenance back to the source chunks.
Then partition that graph with hierarchical Leiden community detection. Leiden finds clusters of densely connected nodes, an "AI ethics" cluster here, a "supply chain" cluster there, and it recurses, so you get a tree: broad communities near the root, tighter sub-topics at the leaves. For every community at every level, an LLM writes a community summary from the entity and relationship descriptions inside it. Those summaries are the payload.
Chunks ──LLM extract──▶ Entities + Relationships ──▶ Knowledge Graph
│
Leiden (hierarchical)
│
┌──────────── communities ────┴───────────┐
C0 (root: broad themes) │
C1 (sub-topics) │
C2 (leaf: tight clusters) ── LLM ▶ community summariesQuery time. For a global question, GraphRAG does not retrieve chunks at all. It fans the question out across the community summaries at a chosen level: map each summary to a partial answer with a helpfulness score, then reduce the scored partials into one response. The map step is why it scales past the context window, since no single call ever sees the whole corpus, only precomputed summaries. For a local, entity-anchored question it does the opposite, walking out from the matched entities to their neighbors, relationships, and source chunks. The graph supplies the join that vector search could not.
The multi-hop example, concretely
"Do any flagged vendors share a board member with a sanctioned company?"
- Vector RAG: embeds the query, retrieves chunks about vendors and sanctions. The board-member bridge lives on a bio page that mentions neither "vendor" nor "sanctioned," so it never surfaces. Miss.
- GraphRAG: the board member is a node. Edges connect that node to both companies, because extraction read both documents. Traverse two hops and the shared node is right there, with citations. Hit.
The graph made the latent join explicit at index time, so query time is a walk instead of a lucky guess.
The honest numbers
The paper evaluates two corpora: podcast transcripts (~1 million tokens) and news articles (~1.7 million tokens). An LLM judge scored answers on comprehensiveness, diversity, empowerment, and directness. Across community levels and both datasets, GraphRAG beat naive vector RAG on comprehensiveness with win rates of roughly 72–83%, and on diversity of roughly 62–82%. Vector RAG usually won on directness, because flat retrieval gives a tighter, more focused answer when the question is narrow. That is the whole trade in one line: graph retrieval buys breadth and connectedness, and it can cost you crispness on simple questions.
The efficiency picture is the reason to care. Answering a global question from the root-level (C0) summaries used 26–33% fewer tokens than summarizing the raw source text, and up to 97% fewer (roughly 9x–43x) than the lowest, most granular community level, while staying competitive on answer quality. The coarse top of the tree is cheap and often good enough for iterative, exploratory questioning; you descend toward the leaves only when a question needs the depth.
When it earns the cost
Indexing is the catch. You run an LLM over every chunk to extract, then again to summarize every community: real money and real hours before a single query lands. Reach for it only when the workload is genuinely global or multi-hop, such as sensemaking over a fixed corpus, thematic analysis, or connect-the-dots investigation. If your traffic is lookup-shaped ("what's the config flag for X"), a good hybrid retriever (see the earlier lesson on dense plus sparse fusion) will match GraphRAG at a fraction of the cost and beat it on directness.
Builder: Don't rebuild the world. Start with vector retrieval, log the questions it whiffs on, and add a graph layer only if the failures cluster around multi-hop and "summarize all of X." Extraction-prompt quality is your ceiling; a sloppy entity schema poisons everything downstream.
Defender: The graph inherits every attack in its sources. Prompt-injected text in one document does not just corrupt one answer. The injection can be extracted as a false relationship, baked into a community summary, and re-served for months. Treat extraction output as untrusted, keep provenance edges, and make every summary auditable back to source.
Researcher: The LLM-judge win rates are relative preferences, not ground-truth accuracy, and the judge shares a lineage with the generator. Comprehensiveness also rewards longer answers, so check that confound before citing the numbers as capability gains.
The mental model to keep: flat RAG retrieves passages, GraphRAG retrieves structure. Most days you want passages.
Sources
- Edge, Trinh, Cheng, Bradley, Chao, Mody, Truitt, Metropolitansky, Ness, Larson. "From Local to Global: A Graph RAG Approach to Query-Focused Summarization." arXiv:2404.16130, 2024. https://arxiv.org/abs/2404.16130
- Microsoft Research publication page: https://www.microsoft.com/en-us/research/publication/from-local-to-global-a-graph-rag-approach-to-query-focused-summarization/
- Microsoft GraphRAG documentation (Global Search): https://microsoft.github.io/graphrag/
- Microsoft Research blog. "GraphRAG: Improving global search via dynamic community selection." https://www.microsoft.com/en-us/research/blog/graphrag-improving-global-search-via-dynamic-community-selection/
- Traag, Waltman, van Eck. "From Louvain to Leiden: guaranteeing well-connected communities." Scientific Reports, 2019. https://www.nature.com/articles/s41598-019-41695-z