English

RAGXplain: From Explainable Evaluation to Actionable Guidance of RAG Pipelines

Information Retrieval 2026-03-19 v2 Artificial Intelligence

Abstract

Retrieval-Augmented Generation (RAG) systems couple large language models with external knowledge, yet most evaluation methods report aggregate scores that reveal whether a pipeline underperforms but not where or why. We introduce RAGXplain, an evaluation framework that translates performance metrics into actionable guidance. RAGXplain structures evaluation around a 'Metric Diamond' connecting user input, retrieved context, generated answer, and (when available) ground truth via six diagnostic dimensions. It uses LLM reasoning to produce natural-language failure-mode explanations and prioritized interventions. Across five QA benchmarks, applying RAGXplain's recommendations in a single human-guided pass consistently improves RAG pipeline performance across multiple metrics. We release RAGXplain as open source to support reproducibility and community adoption.

Keywords

Cite

@article{arxiv.2505.13538,
  title  = {RAGXplain: From Explainable Evaluation to Actionable Guidance of RAG Pipelines},
  author = {Dvir Cohen and Tamir Houri and Lin Burg and Gilad Barkan},
  journal= {arXiv preprint arXiv:2505.13538},
  year   = {2026}
}