English
Related papers

Related papers: Recommendations and Reporting Checklist for Rigoro…

200 papers

Explainable artificial intelligence techniques are developed at breakneck speed, but suitable evaluation approaches lag behind. With explainers becoming increasingly complex and a lack of consensus on how to assess their utility, it is…

Human-Computer Interaction · Computer Science 2023-04-18 Edward Small , Yueqing Xuan , Danula Hettiachchi , Kacper Sokol

Deep neural networks excel in medical imaging but remain prone to biases, leading to fairness gaps across demographic groups. We provide the first systematic exploration of Human-AI alignment and fairness in this domain. Our results show…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Haozhe Luo , Ziyu Zhou , Zixin Shu , Aurélie Pahud de Mortanges , Robert Berke , Mauricio Reyes

This work establishes a foundational framework for standardizing AI evaluation RCTs (sometimes called human uplift studies). Drawing on established experimental practices from disciplines with established RCT traditions, including software…

Leaderboards are crucial in the machine learning (ML) domain for benchmarking and tracking progress. However, creating leaderboards traditionally demands significant manual effort. In recent years, efforts have been made to automate…

Machine Learning · Computer Science 2026-02-02 Roelien C. Timmer , Necva Bölücü , Stephen Wan

There is no consensus on what constitutes human-centeredness in AI, and existing frameworks lack empirical validation. This study addresses this gap by developing a hierarchical framework of 26 attributes of human-centeredness, validated…

Human-Computer Interaction · Computer Science 2025-02-06 Aung Pyae

Two main approaches for evaluating the quality of machine-generated rationales are: 1) using human rationales as a gold standard; and 2) automated metrics based on how rationales affect model behavior. An open question, however, is how…

Computation and Language · Computer Science 2020-10-13 Samuel Carton , Anirudh Rathore , Chenhao Tan

Benchmarking has long served as a foundational practice in machine learning and, increasingly, in modern AI systems such as large language models, where shared tasks, metrics, and leaderboards offer a common basis for measuring progress and…

Artificial Intelligence · Computer Science 2026-02-16 Philip Waggoner

Generative Artificial Intelligence (AI) holds immense potential in medical applications. Numerous studies have explored the efficacy of various generative AI models within healthcare contexts, but there is a lack of a comprehensive and…

Human-Computer Interaction · Computer Science 2023-12-19 Jinghong Chen , Lingxuan Zhu , Weiming Mou , Zaoqu Liu , Quan Cheng , Anqi Lin , Jian Zhang , Peng Luo

Numerical evaluations with comparisons to baselines play a central role when judging research in recommender systems. In this paper, we show that running baselines properly is difficult. We demonstrate this issue on two extensively studied…

Information Retrieval · Computer Science 2019-05-07 Steffen Rendle , Li Zhang , Yehuda Koren

In AI-rich higher education, polished written mathematics has become easier to produce than trustworthy evidence of understanding. This article develops a human-scale methodology for service mathematics, with informatics as its main running…

History and Overview · Mathematics 2026-04-13 Siniša Miličić

AI measurement science has a wide variety of methodologies and measurements for comparing AI systems, resulting in what often appear to be "apples-to-oranges" comparisons across AI evaluations. To move toward "apples-to-apples" comparisons…

Human-Computer Interaction · Computer Science 2026-05-11 Yee-Yin Choong , Kristen Greene , Alice Qian , Meryem Marasli , Ziqi Yang , Sophia Chen , Laura Dabbish , Anand Rao , Hong Shen

In reaction to growing concerns about the potential harms of artificial intelligence (AI), societies have begun to demand more transparency about how AI models and systems are created and used. To address these concerns, several efforts…

Computers and Society · Computer Science 2024-03-13 David Piorkowski , John Richards , Michael Hind

Estimating the performance of a machine learning system is a longstanding challenge in artificial intelligence research. Today, this challenge is especially relevant given the emergence of systems which appear to increasingly outperform…

Machine Learning · Computer Science 2021-09-17 Qiongkai Xu , Christian Walder , Chenchen Xu

AI models are increasingly prevalent in high-stakes environments, necessitating thorough assessment of their capabilities and risks. Benchmarks are popular for measuring these attributes and for comparing model performance, tracking…

Artificial Intelligence · Computer Science 2024-11-21 Anka Reuel , Amelia Hardy , Chandler Smith , Max Lamparth , Malcolm Hardy , Mykel J. Kochenderfer

Advances in foundation modeling have reshaped computational pathology. However, the increasing number of available models and lack of standardized benchmarks make it increasingly complex to assess their strengths, limitations, and potential…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Andrew Zhang , Guillaume Jaume , Anurag Vaidya , Tong Ding , Faisal Mahmood

Machine learning (ML) methods are proliferating in scientific research. However, the adoption of these methods has been accompanied by failures of validity, reproducibility, and generalizability. These failures can hinder scientific…

With the rise of AI systems in real-world applications comes the need for reliable and trustworthy AI. An essential aspect of this are explainable AI systems. However, there is no agreed standard on how explainable AI systems should be…

Artificial Intelligence · Computer Science 2022-07-04 Sascha Saralajew , Ammar Shaker , Zhao Xu , Kiril Gashteovski , Bhushan Kotnis , Wiem Ben Rim , Jürgen Quittek , Carolin Lawrence

As machine learning systems move from theory to practice, they are increasingly tasked with decisions that affect healthcare access, financial opportunities, hiring, and public services. In these contexts, accuracy is only one piece of the…

Machine Learning · Computer Science 2025-11-13 David Sanchez , Holly Lopez , Michelle Buraczyk , Anantaa Kotal

With the availability of large databases and recent improvements in deep learning methodology, the performance of AI systems is reaching or even exceeding the human level on an increasing number of complex tasks. Impressive examples of this…

Artificial Intelligence · Computer Science 2017-08-29 Wojciech Samek , Thomas Wiegand , Klaus-Robert Müller