中文
相关论文

相关论文: A critical analysis of metrics used for measuring …

200 篇论文

While the importance of automatic image analysis is continuously increasing, recent meta-research revealed major flaws with respect to algorithm validation. Performance metrics are particularly key for meaningful, objective, and transparent…

图像与视频处理 · 电气工程与系统科学 2023-12-08 Annika Reinke , Minu D. Tizabi , Carole H. Sudre , Matthias Eisenmann , Tim Rädsch , Michael Baumgartner , Laura Acion , Michela Antonelli , Tal Arbel , Spyridon Bakas , Peter Bankhead , Arriel Benis , Matthew Blaschko , Florian Buettner , M. Jorge Cardoso , Jianxu Chen , Veronika Cheplygina , Evangelia Christodoulou , Beth Cimini , Gary S. Collins , Sandy Engelhardt , Keyvan Farahani , Luciana Ferrer , Adrian Galdran , Bram van Ginneken , Ben Glocker , Patrick Godau , Robert Haase , Fred Hamprecht , Daniel A. Hashimoto , Doreen Heckmann-Nötzel , Peter Hirsch , Michael M. Hoffman , Merel Huisman , Fabian Isensee , Pierre Jannin , Charles E. Kahn , Dagmar Kainmueller , Bernhard Kainz , Alexandros Karargyris , Alan Karthikesalingam , A. Emre Kavur , Hannes Kenngott , Jens Kleesiek , Andreas Kleppe , Sven Kohler , Florian Kofler , Annette Kopp-Schneider , Thijs Kooi , Michal Kozubek , Anna Kreshuk , Tahsin Kurc , Bennett A. Landman , Geert Litjens , Amin Madani , Klaus Maier-Hein , Anne L. Martel , Peter Mattson , Erik Meijering , Bjoern Menze , David Moher , Karel G. M. Moons , Henning Müller , Brennan Nichyporuk , Felix Nickel , M. Alican Noyan , Jens Petersen , Gorkem Polat , Susanne M. Rafelski , Nasir Rajpoot , Mauricio Reyes , Nicola Rieke , Michael Riegler , Hassan Rivaz , Julio Saez-Rodriguez , Clara I. Sánchez , Julien Schroeter , Anindo Saha , M. Alper Selver , Lalith Sharan , Shravya Shetty , Maarten van Smeden , Bram Stieltjes , Ronald M. Summers , Abdel A. Taha , Aleksei Tiulpin , Sotirios A. Tsaftaris , Ben Van Calster , Gaël Varoquaux , Manuel Wiesenfarth , Ziv R. Yaniv , Paul Jäger , Lena Maier-Hein

Recent benchmark studies have claimed that AI has approached or even surpassed human-level performances on various cognitive tasks. However, this position paper argues that current AI evaluation paradigms are insufficient for assessing…

Optimizing a given metric is a central aspect of most current AI approaches, yet overemphasizing metrics leads to manipulation, gaming, a myopic focus on short-term goals, and other unexpected negative consequences. This poses a fundamental…

计算机与社会 · 计算机科学 2020-02-21 Rachel Thomas , David Uminsky

With the rise in high resolution remote sensing technologies there has been an explosion in the amount of data available for forest monitoring, and an accompanying growth in artificial intelligence applications to automatically derive…

Several benchmarks have been built with heavy investment in resources to track our progress in NLP. Thousands of papers published in response to those benchmarks have competed to top leaderboards, with models often surpassing human…

计算与语言 · 计算机科学 2022-10-17 Swaroop Mishra , Anjana Arunkumar , Chris Bryan , Chitta Baral

As frontier artificial intelligence (AI) models rapidly advance, benchmarks are integral to comparing different models and measuring their progress in different task-specific domains. However, there is a lack of guidance on when and how…

计算机与社会 · 计算机科学 2025-07-10 Ayrton San Joaquin , Rokas Gipiškis , Leon Staufer , Ariel Gil

As artificial intelligence increasingly influences our world, it becomes crucial to assess its technical progress and societal impact. This paper surveys problems and opportunities in the measurement of AI systems and their impact, based on…

计算机与社会 · 计算机科学 2020-09-22 Saurabh Mishra , Jack Clark , C. Raymond Perrault

Publicly accessible benchmarks that allow for assessing and comparing model performances are important drivers of progress in artificial intelligence (AI). While recent advances in AI capabilities hold the potential to transform medical…

人工智能 · 计算机科学 2022-12-26 Kathrin Blagec , Jakob Kraiger , Wolfgang Frühwirt , Matthias Samwald

More than one hundred benchmarks have been developed to test the commonsense knowledge and commonsense reasoning abilities of artificial intelligence (AI) systems. However, these benchmarks are often flawed and many aspects of common sense…

人工智能 · 计算机科学 2023-02-24 Ernest Davis

Language models (LMs) represent an emerging paradigm within artificial intelligence, with applications throughout the medical enterprise. A comprehensive understanding of the clinical task and awareness of the variability in performance…

机器学习 · 计算机科学 2026-03-09 Victor Garcia , Mariia Sidulova , Aldo Badano

Context: Software process improvement (SPI) is known as a key for being successfull in software development. Measuring quality and performance is of high importance in agile software development as agile approaches focussing strongly on…

软件工程 · 计算机科学 2024-07-10 Kevin Phong Pham , Michael Neumann

There is a tendency across different subfields in AI to valorize a small collection of influential benchmarks. These benchmarks operate as stand-ins for a range of anointed common problems that are frequently framed as foundational…

机器学习 · 计算机科学 2021-12-01 Inioluwa Deborah Raji , Emily M. Bender , Amandalynne Paullada , Emily Denton , Alex Hanna

Over the Eight decades, computing paradigms have shifted from large, centralized systems to compact, distributed architectures, leading to the rise of the Distributed Computing Continuum (DCC). In this model, multiple layers such as cloud,…

分布式、并行与集群计算 · 计算机科学 2025-12-10 Praveen Kumar Donta , Qiyang Zhang , Schahram Dustdar

Incomplete data are common in practical applications. Most predictive machine learning models do not handle missing values so they require some preprocessing. Although many algorithms are used for data imputation, we do not understand the…

机器学习 · 统计学 2020-07-07 Katarzyna Woźnica , Przemysław Biecek

Algorithmic fairness is receiving significant attention in the academic and broader literature due to the increasing use of predictive algorithms, including those based on artificial intelligence. One benefit of this trend is that algorithm…

计算机与社会 · 计算机科学 2020-01-28 Pratyush Garg , John Villasenor , Virginia Foggo

Existing AI evaluation practices often fail to capture how systems actually perform in low-resource environments, where operational constraints shape usability as much as model quality. Through a structured analysis of existing benchmark…

人工智能 · 计算机科学 2026-05-28 Aakash Pant , Kavya Shah , Apoorv Agnihotri , Sneha Nikam , Prasaanth Balraj , Nakul Jain

Most AI benchmarks saturate within years or even months after they are introduced, making it hard to study long-run trends in AI capabilities. To address this challenge, we build a statistical framework that stitches benchmarks together,…

人工智能 · 计算机科学 2025-12-02 Anson Ho , Jean-Stanislas Denain , David Atanasov , Samuel Albanie , Rohin Shah

Evaluations of generative models are now ubiquitous, and their outcomes critically shape public and scientific expectations of AI's capabilities. Yet skepticism about their reliability continues to grow. How can we know that a reported…

人工智能 · 计算机科学 2026-05-19 Nathanael Jo , Ashia Wilson

The rapid advancement of Artificial Intelligence (AI) has created unprecedented demands for computational power, yet methods for evaluating the performance, efficiency, and environmental impact of deployed models remain fragmented. Current…

性能 · 计算机科学 2025-10-22 Hongyuan Liu , Xinyang Liu , Guosheng Hu

In the rapidly evolving field of artificial intelligence (AI), traditional benchmarks can fall short in attempting to capture the nuanced capabilities of AI models. We focus on the case of physical world modeling and propose a novel…

人工智能 · 计算机科学 2025-09-08 Sasha Mitts