中文
相关论文

相关论文: Who wants accurate models? Arguing for a different…

200 篇论文

Knowing when a classifier's prediction can be trusted is useful in many applications and critical for safely using AI. While the bulk of the effort in machine learning research has been towards improving classifier performance,…

机器学习 · 统计学 2018-10-30 Heinrich Jiang , Been Kim , Melody Y. Guan , Maya Gupta

The H-measure is a classifier performance measure which takes into account the context of application without requiring a rigid value of relative misclassification costs to be set. Since its introduction in 2009 it has become widely…

机器学习 · 计算机科学 2022-01-03 D. J. Hand , C. Anagnostopoulos

An accurate and fair assessment of the efficiency and impact of scientific work is, despite a lot of recent research effort, still an open problem. The measurement of quality and success of individual scientists and research groups can be…

数字图书馆 · 计算机科学 2015-11-19 Miloš Kudělka , Jan Platoš , Pavel Krömer

Evaluating the performance of classifiers is critical in machine learning, particularly in high-stakes applications where the reliability of predictions can significantly impact decision-making. Traditional performance measures, such as…

机器学习 · 计算机科学 2024-12-19 Jesus S. Aguilar-Ruiz

The use of artificial intelligence (AI) in working environments with individuals, known as Human-AI Collaboration (HAIC), has become essential in a variety of domains, boosting decision-making, efficiency, and innovation. Despite HAIC's…

人机交互 · 计算机科学 2025-03-10 George Fragiadakis , Christos Diou , George Kousiouris , Mara Nikolaidou

Online and AI-based symptom checkers are applications that assist medical laypeople in diagnosing their symptoms and determining which course of action to take. When evaluating these tools, previous studies primarily used an approach…

人机交互 · 计算机科学 2025-06-30 Marvin Kopka , Markus A. Feufel

The rapid advancement of Artificial Intelligence (AI) has created unprecedented demands for computational power, yet methods for evaluating the performance, efficiency, and environmental impact of deployed models remain fragmented. Current…

性能 · 计算机科学 2025-10-22 Hongyuan Liu , Xinyang Liu , Guosheng Hu

Artificial intelligence (AI) systems increasingly match or surpass human experts in biomedical signal interpretation. However, their effective integration into clinical practice requires more than high predictive accuracy. Clinicians must…

机器学习 · 计算机科学 2025-10-27 Stefan Kraft , Andreas Theissler , Vera Wienhausen-Wilke , Gjergji Kasneci , Hendrik Lensch

The wide-spread adoption of representation learning technologies in clinical decision making strongly emphasizes the need for characterizing model reliability and enabling rigorous introspection of model behavior. While the former need is…

机器学习 · 计算机科学 2020-05-01 Jayaraman J. Thiagarajan , Prasanna Sattigeri , Deepta Rajan , Bindya Venkatesh

A multitude of explainability methods and associated fidelity performance metrics have been proposed to help better understand how modern AI systems make decisions. However, much of the current work has remained theoretical -- without much…

计算机视觉与模式识别 · 计算机科学 2023-02-01 Julien Colin , Thomas Fel , Remi Cadene , Thomas Serre

Algorithmic fairness is receiving significant attention in the academic and broader literature due to the increasing use of predictive algorithms, including those based on artificial intelligence. One benefit of this trend is that algorithm…

计算机与社会 · 计算机科学 2020-01-28 Pratyush Garg , John Villasenor , Virginia Foggo

Representational similarity metrics are fundamental tools in neuroscience and AI, yet we lack systematic comparisons of their discriminative power across model families. We introduce a quantitative framework to evaluate representational…

机器学习 · 计算机科学 2025-12-10 Jialin Wu , Shreya Saha , Yiqing Bo , Meenakshi Khosla

How do we know if two systems - biological or artificial - process information in a similar way? Similarity measures such as linear regression, Centered Kernel Alignment (CKA), Normalized Bures Similarity (NBS), and angular Procrustes…

神经元与认知 · 定量生物学 2024-12-31 Nathan Cloos , Moufan Li , Markus Siegel , Scott L. Brincat , Earl K. Miller , Guangyu Robert Yang , Christopher J. Cueva

Artificial intelligence (AI) has significantly improved medical screening accuracy, particularly in cancer detection and risk assessment. However, traditional classification metrics often fail to account for imbalanced data, varying…

机器学习 · 计算机科学 2025-10-28 Longfei Wei , Fang Sheng , Jianfei Zhang

AI-powered systems have gained widespread popularity in various domains, including Autonomous Vehicles (AVs). However, ensuring their reliability and safety is challenging due to their complex nature. Conventional test adequacy metrics,…

软件工程 · 计算机科学 2023-11-15 Neelofar Neelofar , Aldeida Aleti

This paper introduces \textit{measurement trees}, a novel class of metrics designed to combine various constructs into an interpretable multi-level representation of a measurand. Unlike conventional metrics that yield single values,…

人工智能 · 计算机科学 2025-10-01 Craig Greenberg , Patrick Hall , Theodore Jensen , Kristen Greene , Razvan Amironesei

AI predictive systems are increasingly embedded in decision making pipelines, shaping high stakes choices once made solely by humans. Yet robust decisions under uncertainty still rely on capabilities that current AI lacks: domain knowledge…

人工智能 · 计算机科学 2025-10-28 Sima Noorani , Shayan Kiyani , George Pappas , Hamed Hassani

As artificial intelligence systems grow more powerful, there has been increasing interest in "AI safety" research to address emerging and future risks. However, the field of AI safety remains poorly defined and inconsistently measured,…

Thanks to the great progress of machine learning in the last years, several Artificial Intelligence (AI) techniques have been increasingly moving from the controlled research laboratory settings to our everyday life. AI is clearly…

人工智能 · 计算机科学 2021-06-07 Tatiana Tommasi , Silvia Bucci , Barbara Caputo , Pietro Asinari

The potential risk of AI systems unintentionally embedding and reproducing bias has attracted the attention of machine learning practitioners and society at large. As policy makers are willing to set the standards of algorithms and AI…

人工智能 · 计算机科学 2020-03-17 Boris Ruf , Chaouki Boutharouite , Marcin Detyniecki