📚 · Ground

GraphRAG

Retrieve from a knowledge graph of entities and relationships instead of flat chunks.

In one line

Turn documents into a graph so the model can answer questions no single chunk covers.

ConceptWhat it is

GraphRAG builds a knowledge graph of entities and relationships extracted from the corpus, then retrieves relevant subgraphs or community summaries instead of, or in addition to, flat text chunks. It exists because flat chunk retrieval struggles with questions that require connecting facts across many documents, like "who worked with whom across these ten reports."

By pre-computing structure and hierarchical summaries, the system can answer both narrow entity questions and broad, corpus-wide themes that naive chunk retrieval cannot reach.

How it worksThe mechanics

An LLM pass over the corpus extracts entities and relationships into a graph, related entities are clustered into communities and summarized, and at query time the system retrieves the relevant entities, their connecting edges, and matching community summaries to assemble context for the final answer.

At a glanceSee it

GraphRAG diagram
GraphRAG diagram 1

GraphRAG routes each question to local entity-anchored search or a global reduce over community summaries — the two retrieval modes a flat pipeline collapses into one step.

GraphRAG diagram 2

Indexing is more than extraction — it resolves duplicate entities and runs Leiden clustering into nested community summaries before any query arrives.

When to use itWhere it fits

  • Corpora rich in entities and relationships, like organizational reports, research literature, or investigative datasets.
  • Questions that require connecting facts across many documents, not just one.
  • Whole-corpus summarization questions like "what are the major themes across all these documents."
  • Domains where relationships between entities are as important as the entities themselves.

When NOT to use itLimits & anti-patterns

  • Simple lookup-style corpora, like a single product manual, where there is no meaningful relationship structure to exploit.
  • Cost-sensitive projects, since graph construction requires an expensive upfront LLM extraction pass over the whole corpus.
  • Fast-changing data, since the graph and community summaries need continual, costly re-indexing.

Trade-offsAdvantages & costs

Advantages
  • Answers cross-document, relationship-heavy questions flat RAG cannot.
  • Community summaries enable good whole-corpus, thematic answers.
  • Graph structure is inspectable and explainable, aiding trust and debugging.
  • Reusable across many query types once built.
Trade-offs & costs
  • Expensive, slow upfront graph construction using LLM extraction passes.
  • Requires ongoing maintenance as source documents change.
  • More infrastructure complexity than a plain vector store.
  • Entity extraction errors can propagate into wrong relationships.

ExampleIn the real world

Microsoft's own GraphRAG research pipeline summarizes news corpora into entity graphs and community reports so it can answer "what are the main narratives in this dataset" rather than just point lookups.

ToolsHow to implement it

  • Microsoft GraphRAGopen-source reference implementation for entity extraction and community summarization.
  • Neo4jgraph database for storing and querying the extracted entity-relationship graph.
  • LlamaIndex Property Graph Indexframework support for building and querying graph-based retrieval.
  • LangGraphorchestration for combining graph traversal with generation steps.

Cost & effortWhat it takes

High upfront cost: the extraction and summarization pass over the full corpus uses significant LLM calls; ongoing storage in a graph database and re-indexing add engineering effort beyond plain vector RAG.

A living map of modern AI — kept current every morning