中文

评估医学大语言模型中提示工程技术的准确性与置信度获取

计算机与社会 2025-06-03 v1 人工智能 计算与语言 机器学习

摘要

本文研究提示工程技术如何影响应用于医学语境的大型语言模型(LLM)的准确性与置信度获取。使用跨多个专业的波斯语执业医师考试分层题库数据集,我们评估了五种LLM——GPT-4o、o3-mini、Llama-3.3-70b、Llama-3.1-8b和DeepSeek-v3——在156种配置下的表现。这些配置在温度设置(0.3、0.7、1.0)、提示风格(思维链、少样本、情感、专家模仿)和置信度量表(1-10、1-100)上有所不同。我们使用AUC-ROC、Brier分数和期望校准误差(ECE)来评估置信度与实际表现之间的一致性。思维链提示提高了准确性,但也导致了过度自信,凸显了校准的必要性。情感提示进一步膨胀了置信度,增加了决策失误的风险。较小模型如Llama-3.1-8b在所有指标上均表现不佳,而专有模型显示出更高的准确性,但仍然缺乏校准的置信度。这些结果表明,提示工程必须同时兼顾准确性和不确定性,才能在高风险医疗任务中有效。

关键词

引用

@article{arxiv.2506.00072,
  title  = {Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs},
  author = {Nariman Naderi and Zahra Atf and Peter R Lewis and Aref Mahjoub far and Seyed Amir Ahmad Safavi-Naini and Ali Soroush},
  journal= {arXiv preprint arXiv:2506.00072},
  year   = {2025}
}

备注

This paper was accepted for presentation at the 7th International Workshop on EXplainable, Trustworthy, and Responsible AI and Multi-Agent Systems (EXTRAAMAS 2025). Workshop website: https://extraamas.ehealth.hevs.ch/index.html