English

Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty

Computation and Language 2026-01-23 v2 Artificial Intelligence

Abstract

Current evaluation of large language models (LLMs) overwhelmingly prioritizes accuracy; however, in real-world and safety-critical applications, the ability to abstain when uncertain is equally vital for trustworthy deployment. We introduce MedAbstain, a unified benchmark and evaluation protocol for abstention in medical multiple-choice question answering (MCQA) -- a discrete-choice setting that generalizes to agentic action selection -- integrating conformal prediction, adversarial question perturbations, and explicit abstention options. Our systematic evaluation of both open- and closed-source LLMs reveals that even state-of-the-art, high-accuracy models often fail to abstain with uncertain. Notably, providing explicit abstention options consistently increases model uncertainty and safer abstention, far more than input perturbations, while scaling model size or advanced prompting brings little improvement. These findings highlight the central role of abstention mechanisms for trustworthy LLM deployment and offer practical guidance for improving safety in high-stakes applications.

Keywords

Cite

@article{arxiv.2601.12471,
  title  = {Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty},
  author = {Sravanthi Machcha and Sushrita Yerra and Sahil Gupta and Aishwarya Sahoo and Sharmin Sultana and Hong Yu and Zonghai Yao},
  journal= {arXiv preprint arXiv:2601.12471},
  year   = {2026}
}

Comments

Equal contribution for the first two authors; To appear in proceedings of the Main Conference of the European Chapter of the Association for Computational Linguistics (EACL) 2026