English

The Reasoning Bottleneck in Graph-RAG: Structured Prompting and Context Compression for Multi-Hop QA

Information Retrieval 2026-03-19 v2 Computation and Language

Abstract

Graph-RAG systems achieve strong multi-hop question answering by indexing documents into knowledge graphs, but strong retrieval does not guarantee strong answers. Evaluating KET-RAG, a leading Graph-RAG system, on three multi-hop QA benchmarks (HotpotQA, MuSiQue, 2WikiMultiHopQA), we find that 77% to 91% of questions have the gold answer in the retrieved context, yet accuracy is only 35% to 78%, and 73% to 84% of errors are reasoning failures. We propose two augmentations: (i) SPARQL chain-of-thought prompting, which decomposes questions into triple-pattern queries aligned with the entity-relationship context, and (ii) graph-walk compression, which compresses the context by ~60% via knowledge-graph traversal with no LLM calls. SPARQL CoT improves accuracy by +2 to +14 pp; graph-walk compression adds +6 pp on average when paired with structured prompting on smaller models. Surprisingly, we show that, with question-type routing, a fully augmented budget open-weight Llama-8B model matches or exceeds the unaugmented Llama-70B baseline on all three benchmarks at ~12x lower cost. A replication on LightRAG confirms that our augmentations transfer across Graph-RAG systems.

Keywords

Cite

@article{arxiv.2603.14045,
  title  = {The Reasoning Bottleneck in Graph-RAG: Structured Prompting and Context Compression for Multi-Hop QA},
  author = {Yasaman Zarrinkia and Venkatesh Srinivasan and Alex Thomo},
  journal= {arXiv preprint arXiv:2603.14045},
  year   = {2026}
}

Comments

11 pages, 2 figures, 9 tables; under review

R2 v1 2026-07-01T11:20:13.833Z