中文

大语言模型诊断不确定性估计立场论文:下一个词概率并非诊断前概率

人工智能 2024-11-08 v1 计算与语言

摘要

大语言模型(LLMs)正被探索用于诊断决策支持,但其估计诊断前概率的能力仍然有限,而诊断前概率对临床决策至关重要。本研究使用结构化电子健康记录数据,在三个诊断任务上评估了两个大语言模型:Mistral-7B 和 Llama3-70B。我们审视了当前提取大语言模型概率估计的三种方法,并揭示了它们的局限性。我们的目的是强调改进大语言模型置信度估计技术的必要性。

关键词

引用

@article{arxiv.2411.04962,
  title  = {Position Paper On Diagnostic Uncertainty Estimation from Large Language Models: Next-Word Probability Is Not Pre-test Probability},
  author = {Yanjun Gao and Skatje Myers and Shan Chen and Dmitriy Dligach and Timothy A Miller and Danielle Bitterman and Guanhua Chen and Anoop Mayampurath and Matthew Churpek and Majid Afshar},
  journal= {arXiv preprint arXiv:2411.04962},
  year   = {2024}
}

备注

Accepted to GenAI4Health Workshop at NeurIPS 2024