English

PAR$^2$-RAG: Planned Active Retrieval and Reasoning for Multi-Hop Question Answering

Artificial Intelligence 2026-04-01 v1

Abstract

Large language models (LLMs) remain brittle on multi-hop question answering (MHQA), where answering requires combining evidence across documents through retrieval and reasoning. Iterative retrieval systems can fail by locking onto an early low-recall trajectory and amplifying downstream errors, while planning-only approaches may produce static query sets that cannot adapt when intermediate evidence changes. We propose \textbf{Planned Active Retrieval and Reasoning RAG (PAR2^2-RAG)}, a two-stage framework that separates \emph{coverage} from \emph{commitment}. PAR2^2-RAG first performs breadth-first anchoring to build a high-recall evidence frontier, then applies depth-first refinement with evidence sufficiency control in an iterative loop. Across four MHQA benchmarks, PAR2^2-RAG consistently outperforms existing state-of-the-art baselines, compared with IRCoT, PAR2^2-RAG achieves up to \textbf{23.5\%} higher accuracy, with retrieval gains of up to \textbf{10.5\%} in NDCG.

Keywords

Cite

@article{arxiv.2603.29085,
  title  = {PAR$^2$-RAG: Planned Active Retrieval and Reasoning for Multi-Hop Question Answering},
  author = {Xingyu Li and Rongguang Wang and Yuying Wang and Mengqing Guo and Chenyang Li and Tao Sheng and Sujith Ravi and Dan Roth},
  journal= {arXiv preprint arXiv:2603.29085},
  year   = {2026}
}

Comments

11 pages, 2 figures