English

Understanding metric-related pitfalls in image analysis validation

Computer Vision and Pattern Recognition 2024-02-26 v4

Abstract

Validation metrics are key for the reliable tracking of scientific progress and for bridging the current chasm between artificial intelligence (AI) research and its translation into practice. However, increasing evidence shows that particularly in image analysis, metrics are often chosen inadequately in relation to the underlying research problem. This could be attributed to a lack of accessibility of metric-related knowledge: While taking into account the individual strengths, weaknesses, and limitations of validation metrics is a critical prerequisite to making educated choices, the relevant knowledge is currently scattered and poorly accessible to individual researchers. Based on a multi-stage Delphi process conducted by a multidisciplinary expert consortium as well as extensive community feedback, the present work provides the first reliable and comprehensive common point of access to information on pitfalls related to validation metrics in image analysis. Focusing on biomedical image analysis but with the potential of transfer to other fields, the addressed pitfalls generalize across application domains and are categorized according to a newly created, domain-agnostic taxonomy. To facilitate comprehension, illustrations and specific examples accompany each pitfall. As a structured body of information accessible to researchers of all levels of expertise, this work enhances global comprehension of a key topic in image analysis validation.

Keywords

Cite

@article{arxiv.2302.01790,
  title  = {Understanding metric-related pitfalls in image analysis validation},
  author = {Annika Reinke and Minu D. Tizabi and Michael Baumgartner and Matthias Eisenmann and Doreen Heckmann-Nötzel and A. Emre Kavur and Tim Rädsch and Carole H. Sudre and Laura Acion and Michela Antonelli and Tal Arbel and Spyridon Bakas and Arriel Benis and Matthew Blaschko and Florian Buettner and M. Jorge Cardoso and Veronika Cheplygina and Jianxu Chen and Evangelia Christodoulou and Beth A. Cimini and Gary S. Collins and Keyvan Farahani and Luciana Ferrer and Adrian Galdran and Bram van Ginneken and Ben Glocker and Patrick Godau and Robert Haase and Daniel A. Hashimoto and Michael M. Hoffman and Merel Huisman and Fabian Isensee and Pierre Jannin and Charles E. Kahn and Dagmar Kainmueller and Bernhard Kainz and Alexandros Karargyris and Alan Karthikesalingam and Hannes Kenngott and Jens Kleesiek and Florian Kofler and Thijs Kooi and Annette Kopp-Schneider and Michal Kozubek and Anna Kreshuk and Tahsin Kurc and Bennett A. Landman and Geert Litjens and Amin Madani and Klaus Maier-Hein and Anne L. Martel and Peter Mattson and Erik Meijering and Bjoern Menze and Karel G. M. Moons and Henning Müller and Brennan Nichyporuk and Felix Nickel and Jens Petersen and Susanne M. Rafelski and Nasir Rajpoot and Mauricio Reyes and Michael A. Riegler and Nicola Rieke and Julio Saez-Rodriguez and Clara I. Sánchez and Shravya Shetty and Maarten van Smeden and Ronald M. Summers and Abdel A. Taha and Aleksei Tiulpin and Sotirios A. Tsaftaris and Ben Van Calster and Gaël Varoquaux and Manuel Wiesenfarth and Ziv R. Yaniv and Paul F. Jäger and Lena Maier-Hein},
  journal= {arXiv preprint arXiv:2302.01790},
  year   = {2024}
}

Comments

Shared first authors: Annika Reinke and Minu D. Tizabi; shared senior authors: Lena Maier-Hein and Paul F. J\"ager. Published in Nature Methods. arXiv admin note: text overlap with arXiv:2206.01653

R2 v1 2026-06-28T08:31:26.145Z