中文

医学影像AI中的虚假承诺?评估性能超越声明的有效性

计算机视觉与模式识别 2025-08-07 v2

摘要

性能比较是医学影像人工智能(AI)研究的基础,通常基于常用性能指标的相对改进来驱动优越性声明。然而,此类声明常常仅依赖于经验平均性能。在本文中,我们通过分析一组具有代表性的医学影像论文,研究新提出的方法是否真正超越了现有技术。我们基于贝叶斯方法量化虚假声明的概率,该方法利用报告的结果以及经验估计的模型一致性来估计方法的相对排名是否可能偶然发生。根据我们的结果,大多数(>80%)论文在引入新方法时声称性能超越。我们的分析进一步揭示,在86%的分类论文和53%的分割论文中,存在高概率(>5%)的虚假性能超越声明。这些发现凸显了当前基准测试实践中的一个关键缺陷:医学影像AI中的性能超越声明常常缺乏充分依据,存在误导未来研究工作的风险。

关键词

引用

@article{arxiv.2505.04720,
  title  = {False Promises in Medical Imaging AI? Assessing Validity of Outperformance Claims},
  author = {Evangelia Christodoulou and Annika Reinke and Pascaline Andrè and Patrick Godau and Piotr Kalinowski and Rola Houhou and Selen Erkan and Carole H. Sudre and Ninon Burgos and Sofiène Boutaj and Sophie Loizillon and Maëlys Solal and Veronika Cheplygina and Charles Heitz and Michal Kozubek and Michela Antonelli and Nicola Rieke and Antoine Gilson and Leon D. Mayer and Minu D. Tizabi and M. Jorge Cardoso and Amber Simpson and Annette Kopp-Schneider and Gaël Varoquaux and Olivier Colliot and Lena Maier-Hein},
  journal= {arXiv preprint arXiv:2505.04720},
  year   = {2025}
}