中文

VeriFact:利用电子健康记录验证LLM生成的临床文本中的事实

人工智能 2025-01-29 v1 计算与语言 信息检索 计算机科学中的逻辑

摘要

目前缺乏确保大型语言模型(LLM)在临床医学中生成的文本事实准确性的方法。VeriFact是一个人工智能系统,它结合了检索增强生成和LLM-as-a-Judge,以验证LLM生成的文本是否基于患者的电子健康记录(EHR)得到事实支持。为了评估该系统,我们引入了VeriFact-BHC,这是一个新数据集,将出院小结中的简要住院病程叙述分解为一组简单陈述,并附有临床医生注释,说明每个陈述是否得到患者EHR临床记录的支持。尽管临床医生之间的最高一致性为88.5%,但与去噪和裁决后的平均人类临床医生真实值相比,VeriFact达到高达92.7%的一致性,表明VeriFact超过了普通临床医生根据患者病历核查文本事实的能力。VeriFact可能通过消除当前的评估瓶颈来加速基于LLM的EHR应用的开发。

关键词

引用

@article{arxiv.2501.16672,
  title  = {VeriFact: Verifying Facts in LLM-Generated Clinical Text with Electronic Health Records},
  author = {Philip Chung and Akshay Swaminathan and Alex J. Goodell and Yeasul Kim and S. Momsen Reincke and Lichy Han and Ben Deverett and Mohammad Amin Sadeghi and Abdel-Badih Ariss and Marc Ghanem and David Seong and Andrew A. Lee and Caitlin E. Coombes and Brad Bradshaw and Mahir A. Sufian and Hyo Jung Hong and Teresa P. Nguyen and Mohammad R. Rasouli and Komal Kamra and Mark A. Burbridge and James C. McAvoy and Roya Saffary and Stephen P. Ma and Dev Dash and James Xie and Ellen Y. Wang and Clifford A. Schmiesing and Nigam Shah and Nima Aghaeepour},
  journal= {arXiv preprint arXiv:2501.16672},
  year   = {2025}
}

备注

62 pages, 5 figures, 1 table, pre-print manuscript