English
Related papers

Related papers: Statistical Confidence in Functional Correctness: …

200 papers

When used in complex engineered systems, such as communication networks, artificial intelligence (AI) models should be not only as accurate as possible, but also well calibrated. A well-calibrated AI model is one that can reliably quantify…

Machine Learning · Computer Science 2022-12-16 Kfir M. Cohen , Sangwoo Park , Osvaldo Simeone , Shlomo Shamai

Constructing valid confidence sets is a crucial task in statistical inference, yet traditional methods often face challenges when dealing with complex models or limited observed sample sizes. These challenges are frequently encountered in…

Thanks to the great progress of machine learning in the last years, several Artificial Intelligence (AI) techniques have been increasingly moving from the controlled research laboratory settings to our everyday life. AI is clearly…

Artificial Intelligence · Computer Science 2021-06-07 Tatiana Tommasi , Silvia Bucci , Barbara Caputo , Pietro Asinari

Software testing is a crucial phase in the software development lifecycle (SDLC), ensuring that products meet necessary functional, performance, and quality benchmarks before release. Despite advancements in automation, traditional methods…

Software Engineering · Computer Science 2026-03-10 Mohammad Baqar , Rajat Khanda

Artificial intelligence (AI) has revolutionized decision-making processes and systems throughout society and, in particular, has emerged as a significant technology in high-impact scenarios of national interest. Yet, despite AI's impressive…

Machine Learning · Statistics 2024-08-05 Gregory Canal , Vladimir Leung , Philip Sage , Eric Heim , I-Jeng Wang

Uncertainty quantification is a set of techniques that measure confidence in language models. They can be used, for example, to detect hallucinations or alert users to review uncertain predictions. To be useful, these confidence scores must…

Computation and Language · Computer Science 2026-04-13 Lorenzo Jaime Yu Flores , Cesare Spinoso di-Piano , Jackie Chi Kit Cheung

For an AI solution to evolve from a trained machine learning model into a production-ready AI system, many more things need to be considered than just the performance of the machine learning model. A production-ready AI system needs to be…

Artificial Intelligence · Computer Science 2023-03-24 Petra Heck , Gerard Schouten

Comprehensive and accurate evaluation of general-purpose AI systems such as large language models allows for effective mitigation of their risks and deepened understanding of their capabilities. Current evaluation methodology, mostly based…

Artificial Intelligence · Computer Science 2024-01-01 Xiting Wang , Liming Jiang , Jose Hernandez-Orallo , David Stillwell , Luning Sun , Fang Luo , Xing Xie

This white paper examines the technical foundations of European AI standardization under the AI Act. It explains how harmonized standards enable the presumption of conformity mechanism, describes the CEN/CENELEC standardization process, and…

Computers and Society · Computer Science 2026-02-03 Piercosma Bisconti , Marcello Galisai

Today, AI is being increasingly used to help human experts make decisions in high-stakes scenarios. In these scenarios, full automation is often undesirable, not only due to the significance of the outcome, but also because human experts…

Artificial Intelligence · Computer Science 2020-01-08 Yunfeng Zhang , Q. Vera Liao , Rachel K. E. Bellamy

In industrial settings, surface defects on steel can significantly compromise its service life and elevate potential safety risks. Traditional defect detection methods predominantly rely on manual inspection, which suffers from low…

Machine Learning · Computer Science 2025-04-25 Cheng Shen , Yuewei Liu

Developing and implementing AI-based solutions help state and federal government agencies, research institutions, and commercial companies enhance decision-making processes, automate chain operations, and reduce the consumption of natural…

Artificial Intelligence · Computer Science 2021-12-02 Andrei Svetovidov , Abdul Rahman , Feras A. Batarseh

Providing well-calibrated AI confidence can help promote users' appropriate trust in and reliance on AI, which are essential for AI-assisted decision-making. However, calibrating AI confidence -- providing confidence score that accurately…

Artificial Intelligence · Computer Science 2025-09-30 Jingshu Li , Yitian Yang , Renwen Zhang , Q. Vera Liao , Tianqi Song , Zhengtao Xu , Yi-chieh Lee

We study how people trade off accuracy when using AI-powered tools in professional versus personal contexts for adoption purposes, the determinants of those trade-offs, and how users cope when AI/apps are unavailable. Because modern AI…

Artificial Intelligence · Computer Science 2026-02-17 Gaston Besanson , Federico Todeschini

Verifying the correctness of Bayesian computation is challenging. This is especially true for complex models that are common in practice, as these require sophisticated model implementations and algorithms. In this paper we introduce…

Methodology · Statistics 2020-10-22 Sean Talts , Michael Betancourt , Daniel Simpson , Aki Vehtari , Andrew Gelman

In AI-assisted decision-making, it is crucial but challenging for humans to achieve appropriate reliance on AI. This paper approaches this problem from a human-centered perspective, "human self-confidence calibration". We begin by proposing…

Human-Computer Interaction · Computer Science 2024-03-15 Shuai Ma , Xinru Wang , Ying Lei , Chuhan Shi , Ming Yin , Xiaojuan Ma

The reliability of survey data is crucial in supply chain decision-making, particularly when evaluating readiness for AI-driven tools such as safety stock optimization systems. However, surveys often attract low-effort or fake responses…

Computers and Society · Computer Science 2026-01-27 Bhubalan Mani

Building trust in AI-based systems is deemed critical for their adoption and appropriate use. Recent research has thus attempted to evaluate how various attributes of these systems affect user trust. However, limitations regarding the…

Human-Computer Interaction · Computer Science 2022-04-29 Michaela Benk , Suzanne Tolmeijer , Florian von Wangenheim , Andrea Ferrario

Classification models play a central role in data-driven decision-making applications such as medical diagnosis, recommendation systems, and risk assessment. Traditional performance metrics, such as accuracy and AUC, focus on overall error…

Machine Learning · Computer Science 2026-04-03 Chen Yang , Zheng Cui , Daniel Zhuoyu Long , Jin Qi , Ruohan Zhan

A key trait of stochastic optimizers is that multiple runs of the same optimizer in attempting to solve the same problem can produce different results. As a result, their performance is evaluated over several repeats, or runs, on the…

Machine Learning · Computer Science 2026-05-18 Moslem Noori , Elisabetta Valiante , Thomas Van Vaerenbergh , Masoud Mohseni , Ignacio Rozada