中文
相关论文

相关论文: On the relationship between Benchmarking, Standard…

200 篇论文

Trust in robots is widely believed to be imperative for the adoption of robots into people's daily lives. It is, therefore, understandable that the literature of the last few decades focuses on measuring how much people trust robots -- and…

机器人学 · 计算机科学 2023-11-22 Patrick Holthaus , Alessandra Rossi

This paper is intended to provide an overview of how the evaluation of standards could be applied to entity resolution, or record linkage. Data quality is of critical importance for many AI applications, and the quality of data,…

计算机与社会 · 计算机科学 2025-08-19 Julia Lane

In this work-in-progress, we investigate the certification of AI systems, focusing on the practical application and limitations of existing certification catalogues in the light of the AI Act by attempting to certify a publicly available AI…

计算机与社会 · 计算机科学 2025-02-19 Gregor Autischer , Kerstin Waxnegger , Dominik Kowald

The problem of human trust in artificial intelligence is one of the most fundamental problems in applied machine learning. Our processes for evaluating AI trustworthiness have substantial ramifications for ML's impact on science, health,…

机器学习 · 计算机科学 2022-02-14 Max W. Shen

This study discusses the interplay between metrics used to measure the explainability of the AI systems and the proposed EU Artificial Intelligence Act. A standardisation process is ongoing: several entities (e.g. ISO) and scholars are…

人工智能 · 计算机科学 2022-01-20 Francesco Sovrano , Salvatore Sapienza , Monica Palmirani , Fabio Vitali

AI tools to support real world decision making must be able to build simulation models that inform their recommendations and render them interpretable. Tools that can automate aspects of modeling practice must complement human expertise,…

人工智能 · 计算机科学 2026-05-29 Sara Metcalf , William Schoenberg

Architectures for quantum computing can only be scaled up when they are accompanied by suitable benchmarking techniques. The document provides a comprehensive overview of the state and recommendations for systematic benchmarking of quantum…

Autonomous robotic systems are complex, hybrid, and often safety-critical; this makes their formal specification and verification uniquely challenging. Though commonly used, testing and simulation alone are insufficient to ensure the…

软件工程 · 计算机科学 2021-01-29 Matt Luckcuck , Marie Farrel , Louise A. Dennis , Michael Fisher

Validity, reliability, and fairness are core ethical principles embedded in classical argument-based assessment validation theory. These principles are also central to the Standards for Educational and Psychological Testing (2014) which…

计算机与社会 · 计算机科学 2024-11-06 Jill Burstein , Geoffrey T. LaFlair

Comprehensive and accurate evaluation of general-purpose AI systems such as large language models allows for effective mitigation of their risks and deepened understanding of their capabilities. Current evaluation methodology, mostly based…

人工智能 · 计算机科学 2024-01-01 Xiting Wang , Liming Jiang , Jose Hernandez-Orallo , David Stillwell , Luning Sun , Fang Luo , Xing Xie

For a general standardized testing algorithm designed to evaluate a specific aspect of a robot's performance, several key expectations are commonly imposed. Beyond accuracy (i.e., closeness to a typically unknown ground-truth reference) and…

机器人学 · 计算机科学 2025-12-22 Bowen Weng , Linda Capito , Guillermo A. Castillo , Dylan Khor

Predictive benchmarking, the evaluation of machine learning models based on predictive performance and competitive ranking, is a central epistemic practice in machine learning research and an increasingly prominent method for scientific…

机器学习 · 计算机科学 2025-10-28 Timo Freiesleben , Sebastian Zezulka

This paper reviews and proposes concerns in adopting, fielding, and maintaining artificial intelligence (AI) systems. While the AI community has made rapid progress, there are challenges in certifying AI systems. Using procedures from…

人工智能 · 计算机科学 2021-11-04 Erik Blasch , Junchi Bin , Zheng Liu

We motivate and outline a programme for a formal theory of measurement of artificial intelligence. We argue that formalising measurement for AI will allow researchers, practitioners, and regulators to: (i) make comparisons between systems…

人工智能 · 计算机科学 2025-07-09 Elija Perrier

The creation of benchmarks to evaluate the safety of Large Language Models is one of the key activities within the trusted AI community. These benchmarks allow models to be compared for different aspects of safety such as toxicity, bias,…

人工智能 · 计算机科学 2025-06-23 Lina Berrayana , Sean Rooney , Luis Garcés-Erice , Ioana Giurgiu

Many decision-making processes have begun to incorporate an AI element, including prison sentence recommendations, college admissions, hiring, and mortgage approval. In all of these cases, AI models are being trained to help human decision…

计算机与社会 · 计算机科学 2019-12-06 Maryam Ashoori , Justin D. Weisz

As AI regulations around the world intensify their focus on system safety, contestability has become a mandatory, yet ill-defined, safeguard. In XAI, "contestability" remains an empty promise: no formal definition exists, no algorithm…

计算机与社会 · 计算机科学 2025-06-03 Catarina Moreira , Anna Palatkina , Dacia Braca , Dylan M. Walsh , Peter J. Leihn , Fang Chen , Nina C. Hubig

In this paper, we present an approach for quantifying the propagated uncertainty of robot systems in an online and data-driven manner. Especially in Human-Robot Collaboration, keeping track of the safety compliance during run time is…

机器人学 · 计算机科学 2023-02-22 Woo-Jeong Baek , Torsten Kröger

Assuring safety for ``AI-based'' systems is one of the current challenges in safety engineering. For automated driving systems, in particular, further assurance challenges result from the open context that the systems need to operate in…

系统与控制 · 电气工程与系统科学 2025-07-29 Marcus Nolte , Nayel Fabian Salem , Olaf Franke , Jan Heckmann , Christoph Höhmann , Georg Stettinger , Markus Maurer

As robots find their way into more and more aspects of everyday life, questions around trust are becoming increasingly important. What does it mean to trust a robot? And how should we think about trust in relationships that involve both…