English

Medchain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive Sequence

Computation and Language 2025-10-13 v2 Artificial Intelligence

Abstract

Clinical decision making (CDM) is a complex, dynamic process crucial to healthcare delivery, yet it remains a significant challenge for artificial intelligence systems. While Large Language Model (LLM)-based agents have been tested on general medical knowledge using licensing exams and knowledge question-answering tasks, their performance in the CDM in real-world scenarios is limited due to the lack of comprehensive testing datasets that mirror actual medical practice. To address this gap, we present MedChain, a dataset of 12,163 clinical cases that covers five key stages of clinical workflow. MedChain distinguishes itself from existing benchmarks with three key features of real-world clinical practice: personalization, interactivity, and sequentiality. Further, to tackle real-world CDM challenges, we also propose MedChain-Agent, an AI system that integrates a feedback mechanism and a MCase-RAG module to learn from previous cases and adapt its responses. MedChain-Agent demonstrates remarkable adaptability in gathering information dynamically and handling sequential clinical tasks, significantly outperforming existing approaches.

Keywords

Cite

@article{arxiv.2412.01605,
  title  = {Medchain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive Sequence},
  author = {Jie Liu and Wenxuan Wang and Zizhan Ma and Guolin Huang and Yihang SU and Kao-Jung Chang and Wenting Chen and Haoliang Li and Linlin Shen and Michael Lyu},
  journal= {arXiv preprint arXiv:2412.01605},
  year   = {2025}
}

Comments

Accepted by NeurIPS25 Spotlight

R2 v1 2026-06-28T20:19:54.646Z