English

Position Paper On Diagnostic Uncertainty Estimation from Large Language Models: Next-Word Probability Is Not Pre-test Probability

Artificial Intelligence 2024-11-08 v1 Computation and Language

Abstract

Large language models (LLMs) are being explored for diagnostic decision support, yet their ability to estimate pre-test probabilities, vital for clinical decision-making, remains limited. This study evaluates two LLMs, Mistral-7B and Llama3-70B, using structured electronic health record data on three diagnosis tasks. We examined three current methods of extracting LLM probability estimations and revealed their limitations. We aim to highlight the need for improved techniques in LLM confidence estimation.

Keywords

Cite

@article{arxiv.2411.04962,
  title  = {Position Paper On Diagnostic Uncertainty Estimation from Large Language Models: Next-Word Probability Is Not Pre-test Probability},
  author = {Yanjun Gao and Skatje Myers and Shan Chen and Dmitriy Dligach and Timothy A Miller and Danielle Bitterman and Guanhua Chen and Anoop Mayampurath and Matthew Churpek and Majid Afshar},
  journal= {arXiv preprint arXiv:2411.04962},
  year   = {2024}
}

Comments

Accepted to GenAI4Health Workshop at NeurIPS 2024

R2 v1 2026-06-28T19:52:03.628Z