English
Related papers

Related papers: General Scales Unlock AI Evaluation with Explanato…

200 papers

OpenAI's o3 achieves a high score of 87.5 % on ARC-AGI, a benchmark proposed to measure intelligence. This raises the question whether systems based on Large Language Models (LLMs), particularly o3, demonstrate intelligence and progress…

Artificial Intelligence · Computer Science 2025-01-14 Rolf Pfister , Hansueli Jud

This paper reports a case study on how explainability requirements were elicited during the development of an AI system for predicting cerebral palsy (CP) risk in infants. Over 18 months, we followed a development team and hospital…

Software Engineering · Computer Science 2026-01-06 Tor Sporsem , Stine Rasdal Finserås , Lars Adde , Inga Strümke

As artificial intelligence (AI) becomes integral to economy and society, communication gaps between developers, users, and stakeholders hinder trust and informed decision-making. High-level AI labels, inspired by frameworks like EU energy…

Artificial Intelligence · Computer Science 2025-01-22 Raphael Fischer , Magdalena Wischnewski , Alexander van der Staay , Katharina Poitz , Christian Janiesch , Thomas Liebig

Recent advances in artificial intelligence (AI) have achieved human-scale speed and accuracy for classification tasks. In turn, these capabilities have made AI a viable replacement for many human activities that at their core involve…

Artificial Intelligence · Computer Science 2022-05-24 Hadi Esmaeilzadeh , Reza Vaezi

The integration of artificial intelligence into business processes has significantly enhanced decision-making capabilities across various industries such as finance, healthcare, and retail. However, explaining the decisions made by these AI…

Artificial Intelligence · Computer Science 2024-10-29 Arne Grobrugge , Nidhi Mishra , Johannes Jakubik , Gerhard Satzger

In this paper, we develop the position that current frameworks for evaluating emotional intelligence (EI) in artificial intelligence (AI) systems need refinement because they do not adequately or comprehensively measure the various aspects…

Artificial Intelligence · Computer Science 2025-12-30 Max Parks , Kheli Atluru , Meera Vinod , Mike Kuniavsky , Jud Brewer , Sean White , Sarah Adler , Wendy Ju

As frontier artificial intelligence (AI) models rapidly advance, benchmarks are integral to comparing different models and measuring their progress in different task-specific domains. However, there is a lack of guidance on when and how…

Computers and Society · Computer Science 2025-07-10 Ayrton San Joaquin , Rokas Gipiškis , Leon Staufer , Ariel Gil

While games have been used extensively as milestones to evaluate game-playing AI, there exists no standardised framework for reporting the obtained observations. As a result, it remains difficult to draw general conclusions about the…

Artificial Intelligence · Computer Science 2020-07-07 Vanessa Volz , Boris Naujoks

Recent developments in Generative Artificial Intelligence (GenAI) have created significant uncertainty in education, particularly in terms of assessment practices. Against this backdrop, we present an updated version of the AI Assessment…

Computers and Society · Computer Science 2025-10-01 Mike Perkins , Jasper Roe , Leon Furze

Artificial Intelligence (AI), particularly through the advent of large-scale generative AI (GenAI) models such as Large Language Models (LLMs), has become a transformative element in contemporary technology. While these models have unlocked…

Software Engineering · Computer Science 2024-01-19 Boming Xia , Qinghua Lu , Liming Zhu , Sung Une Lee , Yue Liu , Zhenchang Xing

The rapid adoption of AI agents across domains has made systematic evaluation crucial for ensuring their usefulness and successful production deployment. Evaluation of AI agents typically involves using a fixed set of benchmarks and…

The field of AI research is advancing at an unprecedented pace, enabling automated hypothesis generation and experimental design across diverse domains such as biology, mathematics, and artificial intelligence. Despite these advancements,…

Machine Learning · Computer Science 2025-10-07 Yaowenqi Liu , Bingxu Meng , Rui Pan , Yuxing Liu , Jerry Huang , Jiaxuan You , Tong Zhang

Generative Artificial Intelligence (AI) holds immense potential in medical applications. Numerous studies have explored the efficacy of various generative AI models within healthcare contexts, but there is a lack of a comprehensive and…

Human-Computer Interaction · Computer Science 2023-12-19 Jinghong Chen , Lingxuan Zhu , Weiming Mou , Zaoqu Liu , Quan Cheng , Anqi Lin , Jian Zhang , Peng Luo

A myriad of measures to illustrate performance of predictive artificial intelligence (AI) models have been proposed in the literature. Selecting appropriate performance measures is essential for predictive AI models that are developed to be…

As frontier AI systems advance toward transformative capabilities, we need a parallel transformation in how we measure and evaluate these systems to ensure safety and inform governance. While benchmarks have been the primary method for…

Artificial Intelligence · Computer Science 2025-05-12 Markov Grey , Charbel-Raphaël Segerie

Artificial intelligence (AI) systems are deployed as collaborators in human decision-making. Yet, evaluation practices focus primarily on model accuracy rather than whether human-AI teams are prepared to collaborate safely and effectively.…

Human-Computer Interaction · Computer Science 2026-03-20 Min Hun Lee

Despite widespread discussion of AGI, there is no clear framework for measuring progress toward it. This ambiguity fuels subjective claims, makes it difficult to track progress, and risks hindering responsible governance. As a starting…

Despite conflicting definitions and conceptions of fairness, AI fairness researchers broadly agree that fairness is context-specific. However, when faced with general-purpose AI, which by definition serves a range of contexts, how should we…

Computers and Society · Computer Science 2025-10-08 Vyoma Raman , Judy Hanwen Shen , Andy K. Zhang , Lindsey Gailmard , Rishi Bommasani , Daniel E. Ho , Angelina Wang

Explainability in AI and ML models is critical for fostering trust, ensuring accountability, and enabling informed decision making in high stakes domains. Yet this objective is often unmet in practice. This paper proposes a general purpose…

Statistical Finance · Quantitative Finance 2025-09-03 N. Jean , G. Le Pera

Across academia, industry, and government, there is an increasing awareness that the measurement tasks involved in evaluating generative AI (GenAI) systems are especially difficult. We argue that these measurement tasks are highly…