中文

面向生成式医学图像评估的特征提取:反对一种演进趋势的新证据

计算机视觉与模式识别 2024-10-23 v5

摘要

Fréchet Inception Distance (FID) 是一种广泛使用的合成图像质量评估指标。它依赖于基于 ImageNet 的特征提取器,这使其对医学影像的适用性尚不明确。近期的一种趋势是通过在医学图像上训练的特征提取器来将 FID 适配到医学影像。我们的研究挑战了这一做法,证明了基于 ImageNet 的提取器比其 RadImageNet 对应物更一致且与人类判断更吻合。我们评估了四个医学影像模态和四种数据增强技术下的十六个 StyleGAN2 网络,使用十一个 ImageNet 或 RadImageNet 训练的特征提取器计算 Fréchet 距离(FDs)。通过视觉图灵测试与人类判断的比较表明,基于 ImageNet 的提取器产生的排序与人类判断一致,且由 ImageNet 训练的 SwAV 提取器导出的 FD 与专家评估显著相关。相比之下,基于 RadImageNet 的排序不稳定且与人类判断不一致。我们的发现挑战了普遍假设,提供了新证据:医学图像训练的特征提取器并不会固有地改善 FDs,甚至可能损害其可靠性。我们的代码可在 https://github.com/mckellwoodland/fid-med-eval 获取。

关键词

引用

@article{arxiv.2311.13717,
  title  = {Feature Extraction for Generative Medical Imaging Evaluation: New Evidence Against an Evolving Trend},
  author = {McKell Woodland and Austin Castelo and Mais Al Taie and Jessica Albuquerque Marques Silva and Mohamed Eltaher and Frank Mohn and Alexander Shieh and Suprateek Kundu and Joshua P. Yung and Ankit B. Patel and Kristy K. Brock},
  journal= {arXiv preprint arXiv:2311.13717},
  year   = {2024}
}

备注

This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution is published in LNCS vol. 15012, and is available online at https://doi.org/10.1007/978-3-031-72390-2_9