Skip to main navigation Skip to search Skip to main content

Understanding metric-related pitfalls in image analysis validation

  • Annika Reinke
  • , Minu D. Tizabi
  • , Michael Baumgartner
  • , Matthias Eisenmann
  • , Doreen Heckmann-Nötzel
  • , A. Emre Kavur
  • , Tim Rädsch
  • , Carole H. Sudre
  • , Laura Acion
  • , Michela Antonelli
  • , Tal Arbel
  • , Spyridon Bakas
  • , Arriel Benis
  • , Florian Buettner
  • , M. Jorge Cardoso
  • , Veronika Cheplygina
  • , Jianxu Chen
  • , Evangelia Christodoulou
  • , Beth A. Cimini
  • , Keyvan Farahani
  • Luciana Ferrer, Adrian Galdran, Bram van Ginneken, Ben Glocker, Patrick Godau, Daniel A. Hashimoto, Michael M. Hoffman, Merel Huisman, Fabian Isensee, Pierre Jannin, Charles E. Kahn, Dagmar Kainmueller, Bernhard Kainz, Alexandros Karargyris, Jens Kleesiek, Florian Kofler, Thijs Kooi, Annette Kopp-Schneider, Michal Kozubek, Anna Kreshuk, Tahsin Kurc, Bennett A. Landman, Geert Litjens, Amin Madani, Klaus Maier-Hein, Anne L. Martel, Erik Meijering, Bjoern Menze, Karel G.M. Moons, Henning Müller, Brennan Nichyporuk, Felix Nickel, Jens Petersen, Susanne M. Rafelski, Nasir Rajpoot, Mauricio Reyes, Michael A. Riegler, Nicola Rieke, Julio Saez-Rodriguez, Clara I. Sánchez, Shravya Shetty, Ronald M. Summers, Abdel A. Taha, Aleksei Tiulpin, Sotirios A. Tsaftaris, Ben Van Calster, Gaël Varoquaux, Ziv R. Yaniv, Paul F. Jäger, Lena Maier-Hein
  • German Cancer Research Center
  • Heidelberg University
  • Universitätsklinikum Heidelberg
  • University College London (UCL)
  • King's College London
  • Universidad de Buenos Aires
  • University College London
  • McGill University
  • Indiana University School of Medicine
  • University of Pennsylvania
  • Holon Institute of Technology
  • European Federation for Medical Informatics
  • Johann Wolfgang Goethe University
  • Frankfurt Cancer Institute
  • IT University of Copenhagen
  • Leibniz-Institut für Analytische Wissenschaften
  • The Broad Institute of MIT and Harvard
  • National Cancer Institute (NCI)
  • Ciudad Autónoma de Buenos Aires
  • Pompeu Fabra University (UPF)
  • University of Adelaide
  • Fraunhofer MEVIS
  • Amalia Children's Hospital
  • Imperial College London
  • Princess Margaret Hospital
  • University of Toronto Faculty of Medicine
  • University of Toronto
  • Vector Institute
  • Université de Rennes 1
  • INSERM U70
  • Max Delbrück Center for Molecular Medicine
  • University of Potsdam
  • Friedrich-Alexander Universitat Erlangen-Nurnberg (FAU)
  • IHU Strasbourg
  • University Medicine Essen
  • Helmholtz AI
  • Lunit Inc.
  • Masaryk University
  • European Molecular Biology Laboratory Heidelberg
  • SUNY
  • Vanderbilt University School of Engineering
  • University Health Network
  • Sunnybrook Research Institute
  • University of New South Wales
  • University of Zurich
  • University Medical Center Utrecht
  • University of Applied Sciences Western Switzerland
  • Faculty of Medicine
  • Quebec Artificial Intelligence Institute
  • Universitätsklinikum Hamburg-Eppendorf
  • Allen Institute for Cell Science
  • University of Warwick
  • University of Bern, Faculty of Medicine
  • Inselspital Universitatsspital
  • Simula Metropolitan Center for Digital Engineering
  • UIT The Arctic University of Norway
  • NVIDIA GmbH
  • University of Amsterdam
  • Google Inc
  • National Institutes of Health Clinical Center (NIH)
  • Technische Universität Wien
  • Faculty of Medicine
  • Oulu University Hospital
  • University of Edinburgh
  • Katholieke Universiteit Leuven
  • Leiden University Medical Centre
  • INRIA Saclay
  • National Institute of Allergy and Infectious Diseases (NIAID)

Research output: Contribution to journalArticlepeer-review

131 Scopus citations

Abstract

Validation metrics are key for tracking scientific progress and bridging the current chasm between artificial intelligence research and its translation into practice. However, increasing evidence shows that, particularly in image analysis, metrics are often chosen inadequately. Although taking into account the individual strengths, weaknesses and limitations of validation metrics is a critical prerequisite to making educated choices, the relevant knowledge is currently scattered and poorly accessible to individual researchers. Based on a multistage Delphi process conducted by a multidisciplinary expert consortium as well as extensive community feedback, the present work provides a reliable and comprehensive common point of access to information on pitfalls related to validation metrics in image analysis. Although focused on biomedical image analysis, the addressed pitfalls generalize across application domains and are categorized according to a newly created, domain-agnostic taxonomy. The work serves to enhance global comprehension of a key topic in image analysis validation.

Original languageEnglish
Pages (from-to)182-194
Number of pages13
JournalNature Methods
Volume21
Issue number2
DOIs
StatePublished - Feb 2024
Externally publishedYes

Fingerprint

Dive into the research topics of 'Understanding metric-related pitfalls in image analysis validation'. Together they form a unique fingerprint.

Cite this