English
Related papers

Related papers: Position: AI Evaluation Should Learn from How We T…

200 papers

Learning difficulties pose significant challenges for students, impacting their academic performance and overall educational experience. These difficulties could sometimes put students into a downward spiral that lack of educational…

Human-Computer Interaction · Computer Science 2024-03-12 Aaron Hu

As frontier AI systems advance toward transformative capabilities, we need a parallel transformation in how we measure and evaluate these systems to ensure safety and inform governance. While benchmarks have been the primary method for…

Artificial Intelligence · Computer Science 2025-05-12 Markov Grey , Charbel-Raphaël Segerie

This survey paper chronicles the evolution of evaluation in multimodal artificial intelligence (AI), framing it as a progression of increasingly sophisticated "cognitive examinations." We argue that the field is undergoing a paradigm shift,…

Artificial Intelligence · Computer Science 2026-01-07 Mayank Ravishankara , Varindra V. Persad Maharaj

AI Scaling has traditionally been synonymous with Scaling Up, which builds larger and more powerful models. However, the growing demand for efficiency, adaptability, and collaboration across diverse applications necessitates a broader…

Machine Learning · Computer Science 2025-05-14 Yunke Wang , Yanxi Li , Chang Xu

This paper presents a theoretical framework for addressing the challenges posed by generative artificial intelligence (AI) in higher education assessment through a machine-versus-machine approach. Large language models like GPT-4, Claude,…

Computers and Society · Computer Science 2025-06-04 Mohammad Saleh Torkestani , Taha Mansouri

In recent years, Artificial Intelligence (AI) algorithms have been proven to outperform traditional statistical methods in terms of predictivity, especially when a large amount of data was available. Nevertheless, the "black box" nature of…

Machine Learning · Statistics 2021-10-14 Nicola Picchiotti , Marco Gori

Automated approaches to answer patient-posed health questions are rising, but selecting among systems requires reliable evaluation. The current gold standard for evaluating the free-text artificial intelligence (AI) responses--human expert…

Artificial Intelligence · Computer Science 2026-05-11 Sarvesh Soni , Dina Demner-Fushman

A multitude of explainability methods and associated fidelity performance metrics have been proposed to help better understand how modern AI systems make decisions. However, much of the current work has remained theoretical -- without much…

Computer Vision and Pattern Recognition · Computer Science 2023-02-01 Julien Colin , Thomas Fel , Remi Cadene , Thomas Serre

Current AI evaluation methods, which rely on static, model-only tests, fail to account for harms that emerge through sustained human-AI interaction. As AI systems proliferate and are increasingly integrated into real-world applications,…

Computers and Society · Computer Science 2025-07-31 Lujain Ibrahim , Saffron Huang , Umang Bhatt , Lama Ahmad , Markus Anderljung

Human values and their measurement are long-standing interdisciplinary inquiry. Recent advances in AI have sparked renewed interest in this area, with large language models (LLMs) emerging as both tools and subjects of value measurement.…

Computation and Language · Computer Science 2025-03-07 Haoran Ye , Yuhang Xie , Yuanyi Ren , Hanjun Fang , Xin Zhang , Guojie Song

Human intelligence exhibits a remarkable capacity for rapid adaptation and effective problem-solving in novel and unfamiliar contexts. We argue that this profound adaptability is fundamentally linked to the efficient construction and…

As industry reports claim agentic AI systems deliver double-digit productivity gains and multi-trillion dollar economic potential, the validity of these claims has become critical for investment decisions, regulatory policy, and responsible…

Computers and Society · Computer Science 2025-10-03 Kiana Jafari Meimandi , Gabriela Aránguiz-Dias , Grace Ra Kim , Lana Saadeddin , Allie Griffith , Mykel J. Kochenderfer

Automated hiring systems are among the fastest-developing of all high-stakes AI systems. Among these are algorithmic personality tests that use insights from psychometric testing, and promise to surface personality traits indicative of…

Computers and Society · Computer Science 2022-04-13 Alene K. Rhea , Kelsey Markey , Lauren D'Arinzo , Hilke Schellmann , Mona Sloane , Paul Squires , Falaah Arif Kahn , Julia Stoyanovich

Scientists, policy-makers, business leaders, and members of the public care about what modern artificial intelligence systems are disposed to do. Yet terms such as capabilities, propensities, skills, values, and abilities are routinely used…

Computers and Society · Computer Science 2026-03-03 Konstantinos Voudouris , Mirko Thalmann , Alex Kipnis , José Hernández-Orallo , Eric Schulz

With the rise of Large Language Models (LLMs), AI assistants' ability to utilize tools, especially through API calls, has advanced notably. This progress has necessitated more accurate evaluation methods. Many existing studies adopt static…

Computation and Language · Computer Science 2024-03-28 Honglin Mu , Yang Xu , Yunlong Feng , Xiaofeng Han , Yitong Li , Yutai Hou , Wanxiang Che

Detecting biases in artificial intelligence has become difficult because of the impenetrable nature of deep learning. The central difficulty is in relating unobservable phenomena deep inside models with observable, outside quantities that…

Computation and Language · Computer Science 2019-12-24 Lizhen Liang , Daniel E. Acuna

Every AI benchmark operationalizes theoretical assumptions about the capability it claims to assess. When assumptions function as unexamined commitments, benchmarks stabilize the dominant paradigm by narrowing what counts as progress. Over…

Artificial Intelligence · Computer Science 2026-05-15 Theodore J Kalaitzidis

Artificial Intelligence (AI) is advancing at an unprecedented pace, with clear potential to enhance decision-making and productivity. Yet, the collaborative decision-making process between humans and AI remains underdeveloped, often falling…

Human-Computer Interaction · Computer Science 2025-04-10 Bowen Lou , Tian Lu , T. S. Raghu , Yingjie Zhang

This position paper argues that job exposure to AI should be measured with grounded, evidence-based methods, not inferred from LLM priors alone. Current theoretical exposure measures use zero-shot prompting to classify task-level AI…

Information Retrieval · Computer Science 2026-05-18 Luca Mouchel , Pierre Bouquet , Yossi Sheffi

Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure or superficial variation. We propose a framework to measure consistent behavioral tendencies…

Artificial Intelligence · Computer Science 2026-05-12 Alexandra Yost , Shreyans Jain , Shivam Raval , Grant Corser , Allen Roush , Nina Xu , Jacqueline Hammack , Ravid Shwartz-Ziv , Amirali Abdullah