English

GAPS: A Clinically Grounded, Automated Benchmark for Evaluating AI Clinicians

Computation and Language 2025-12-18 v2

Abstract

Current benchmarks for AI clinician systems, often based on multiple-choice exams or manual rubrics, fail to capture the depth, robustness, and safety required for real-world clinical practice. To address this, we introduce the GAPS framework, a multidimensional paradigm for evaluating Grounding (cognitive depth), Adequacy (answer completeness), Perturbation (robustness), and Safety. Critically, we developed a fully automated, guideline-anchored pipeline to construct a GAPS-aligned benchmark end-to-end, overcoming the scalability and subjectivity limitations of prior work. Our pipeline assembles an evidence neighborhood, creates dual graph and tree representations, and automatically generates questions across G-levels. Rubrics are synthesized by a DeepResearch agent that mimics GRADE-consistent, PICO-driven evidence review in a ReAct loop. Scoring is performed by an ensemble of large language model (LLM) judges. Validation confirmed our automated questions are high-quality and align with clinician judgment (90% agreement, Cohen's Kappa 0.77). Evaluating state-of-the-art models on the benchmark revealed key failure modes: performance degrades sharply with increased reasoning depth (G-axis), models struggle with answer completeness (A-axis), and they are highly vulnerable to adversarial perturbations (P-axis) as well as certain safety issues (S-axis). This automated, clinically-grounded approach provides a reproducible and scalable method for rigorously evaluating AI clinician systems and guiding their development toward safer, more reliable clinical practice. The benchmark dataset GAPS-NSCLC-preview and evaluation code are publicly available at https://huggingface.co/datasets/AQ-MedAI/GAPS-NSCLC-preview and https://github.com/AQ-MedAI/MedicalAiBenchEval.

Keywords

Cite

@article{arxiv.2510.13734,
  title  = {GAPS: A Clinically Grounded, Automated Benchmark for Evaluating AI Clinicians},
  author = {Xiuyuan Chen and Tao Sun and Dexin Su and Ailing Yu and Junwei Liu and Zhe Chen and Gangzeng Jin and Xin Wang and Jingnan Liu and Hansong Xiao and Hualei Zhou and Dongjie Tao and Chunxiao Guo and Minghui Yang and Yuan Xia and Jing Zhao and Qianrui Fan and Yanyun Wang and Shuai Zhen and Kezhong Chen and Jun Wang and Zewen Sun and Heng Zhao and Tian Guan and Shaodong Wang and Geyun Chang and Jiaming Deng and Hongchengcheng Chen and Kexin Feng and Ruzhen Li and Jiayi Geng and Changtai Zhao and Jun Wang and Guihu Lin and Peihao Li and Liqi Liu and Peng Wei and Jian Wang and Jinjie Gu and Ping Wang and Fan Yang},
  journal= {arXiv preprint arXiv:2510.13734},
  year   = {2025}
}
R2 v1 2026-07-01T06:39:20.889Z