Day 4: GraphRAG Explained: Combining Knowledge Graphs with Vector Search

Production RAG Masterclass · Day 4 of 7

Some questions are not “find the nearest paragraph.” They are “walk a path across three documents and explain the join.” Flat vector search was never designed for that.

GraphRAG architecture combining a knowledge graph with vector search

Figure 1. Vectors find seeds. Graphs find relationships.

Thesis Cosine similarity cannot infer multi-hop entity relationships across separate documents. Pair vectors with an entity–relation graph.

Where flat vector RAG breaks

Ask: “Which vendors touch the same data pipeline as Team Atlas, and who owns the SLA?”

The facts live in an org chart, a vendor contract, and a runbook. No single chunk holds the full path. Similarity returns locally relevant paragraphs — not the chain.

Build the graph at ingest

Extract triples while you chunk:

(Entity) -[RELATION]-> (Entity)

(Team Atlas)-[OWNS]->(Billing Pipeline)
(Billing Pipeline)-[DEPENDS_ON]->(Vendor Stripe)
(Vendor Stripe)-[HAS_SLA]->(99.95%)

Store nodes and edges in Neo4j (or similar). Keep chunk text and embeddings in the vector index. Link chunk IDs as node properties so you can jump graph → evidence.

Entity-relation knowledge graph with a highlighted multi-hop path

Figure 2. Local walks answer entity questions; global communities summarize the corpus.

Hybrid Graph + Vector flow

  1. Vector hit — seed entities or chunks close to the query.
  2. Graph expand — walk 1–3 hops from those entities.
  3. Gather evidence — pull linked chunks for the subgraph.
  4. Generate — answer grounded in both path and text.
seeds = vector_search(query, k=8)
entities = extract_entities(seeds)
subgraph = neo4j.run("""
  MATCH path=(e:Entity)-[*1..2]-(n)
  WHERE e.id IN $ids
  RETURN path
""", ids=entities)
context = chunks_for(subgraph) + seeds
answer = llm(query, context)

Local vs global strategies

Local — entity-centric Q&A, incident tracing, “who depends on X?” Neighborhood walks around seed nodes.

Global — corpus-level themes and executive summaries via community detection when the question is about the whole knowledge base.

When GraphRAG is worth the cost

  • Clear entities: people, systems, vendors, policies.
  • Questions that join facts across documents.
  • Compliance, lineage, and ownership queries.

Skip it for FAQ-style single-doc answers — hybrid search from Day 2 is enough.

Key takeaways

  1. Vectors find neighbors; graphs find relationships.
  2. Extract (Entity)-[REL]->(Entity) at ingest.
  3. Hybrid: vector seeds → graph expand → grounded generation.
  4. Local walks for entity Q&A; global communities for corpus summaries.

Series: Production RAG Masterclass
Previous: Day 3 — Query Optimization
Next: Day 5 — Agentic RAG

Comments

Popular posts from this blog

LangChain 2.0 Tutorial: Build an AI Agent with Tools (2025 Edition)

Best AI Tools for Business Owners in 2025: Your Secret Weapon for Super Productivity & More Free Time!

Unlocking the Future: 10 Key Insights into Web3 Technologies