English

Large language models improve physician accuracy but lead to false reliance

Artificial Intelligence 2026-08-01 v1 Applications

Abstract

Retrieval-augmented large language models (LLMs) promise source-linked clinical support, but their value depends on whether displayed evidence guides rather than distorts physician reliance. We developed CORA, an agentic retrieval-augmented LLM, to investigate how source-linked assistance affects physician decision-making. CORA maintained benchmark performance and achieved larger gains on cases published after the models' training-data cutoffs. In a study of 46 physicians, accuracy increased from 70.8% unaided to 82.6% with CORA. Supporting citations predicted correct answers (87.7% vs 65.5%), but citations created an important asymmetry: perceived support increased adoption of correct advice from 34% to 76.9% but when an incorrect LLM answer appeared citation-supported, physician resistance to it fell from 92% to 34.8%. These findings show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk.

Keywords

Cite

@article{arxiv.2608.00817,
  title  = {Large language models improve physician accuracy but lead to false reliance},
  author = {Tirtha Chanda and Christoph Wies and Franziska Schramm and Carina Nogueira Garcia and Nicolas B. Merl and Martin J. Hetz and Jochen S. Utikal and Phillip Tschandl and Cristian Navarrete-Dechent and Alexander Thiem and Jakob N. Kather and Consortium and Titus J. Brinker},
  journal= {arXiv preprint arXiv:2608.00817},
  year   = {2026}
}