English
Related papers

Related papers: Models That Know How Evaluations Are Designed Scor…

200 papers

AI incident reporting requirements are emerging in regulation and policy, yet no operational criteria exist for determining when a detected AI incident warrants escalation beyond national handling to international coordination. This paper…

Computers and Society · Computer Science 2026-05-20 Francesca Gomez , Matthew Ball , Michael Harre , Lydia Preston , Josephine Schwab , Caio Machado

Both humans and machines learn the meaning of unknown words through contextual information in a sentence, but not all contexts are equally helpful for learning. We introduce an effective method for capturing the level of contextual…

Computation and Language · Computer Science 2023-11-10 Sungjin Nam , David Jurgens , Gwen Frishkoff , Kevyn Collins-Thompson

The rapid expansion of large language model (LLM) safety evaluation has produced a substantial benchmark ecosystem, but not a correspondingly coherent measurement ecosystem. We present AISafetyBenchExplorer, a structured catalogue of 195 AI…

Artificial Intelligence · Computer Science 2026-04-24 Abiodun A. Solanke

Internet memes are a powerful form of online communication, yet their nature and reliance on commonsense knowledge make toxicity detection challenging. Identifying key features for meme interpretation and understanding, is a crucial task.…

Computation and Language · Computer Science 2026-03-05 Stefano De Giorgis , Ting-Chih Chen , Filip Ilievski

This study explores whether labeling AI as "trustworthy" or "reliable" influences user perceptions and acceptance of automotive AI technologies. Using a one-way between-subjects design, the research involved 478 online participants who were…

Human-Computer Interaction · Computer Science 2025-07-15 John Dorsch , Ophelia Deroy

The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehensive industry-standard benchmark for assessing AI-product…

Computers and Society · Computer Science 2025-04-22 Shaona Ghosh , Heather Frase , Adina Williams , Sarah Luger , Paul Röttger , Fazl Barez , Sean McGregor , Kenneth Fricklas , Mala Kumar , Quentin Feuillade--Montixi , Kurt Bollacker , Felix Friedrich , Ryan Tsang , Bertie Vidgen , Alicia Parrish , Chris Knotz , Eleonora Presani , Jonathan Bennion , Marisa Ferrara Boston , Mike Kuniavsky , Wiebke Hutiri , James Ezick , Malek Ben Salem , Rajat Sahay , Sujata Goswami , Usman Gohar , Ben Huang , Supheakmungkol Sarin , Elie Alhajjar , Canyu Chen , Roman Eng , Kashyap Ramanandula Manjusha , Virendra Mehta , Eileen Long , Murali Emani , Natan Vidra , Benjamin Rukundo , Abolfazl Shahbazi , Kongtao Chen , Rajat Ghosh , Vithursan Thangarasa , Pierre Peigné , Abhinav Singh , Max Bartolo , Satyapriya Krishna , Mubashara Akhtar , Rafael Gold , Cody Coleman , Luis Oala , Vassil Tashev , Joseph Marvin Imperial , Amy Russ , Sasidhar Kunapuli , Nicolas Miailhe , Julien Delaunay , Bhaktipriya Radharapu , Rajat Shinde , Tuesday , Debojyoti Dutta , Declan Grabb , Ananya Gangavarapu , Saurav Sahay , Agasthya Gangavarapu , Patrick Schramowski , Stephen Singam , Tom David , Xudong Han , Priyanka Mary Mammen , Tarunima Prabhakar , Venelin Kovatchev , Rebecca Weiss , Ahmed Ahmed , Kelvin N. Manyeki , Sandeep Madireddy , Foutse Khomh , Fedor Zhdanov , Joachim Baumann , Nina Vasan , Xianjun Yang , Carlos Mougn , Jibin Rajan Varghese , Hussain Chinoy , Seshakrishna Jitendar , Manil Maskey , Claire V. Hardgrove , Tianhao Li , Aakash Gupta , Emil Joswin , Yifan Mai , Shachi H Kumar , Cigdem Patlak , Kevin Lu , Vincent Alessi , Sree Bhargavi Balija , Chenhe Gu , Robert Sullivan , James Gealy , Matt Lavrisa , James Goel , Peter Mattson , Percy Liang , Joaquin Vanschoren

As AI systems advance beyond human capabilities, scalable oversight becomes critical: how can we supervise AI that exceeds our abilities? A key challenge is that human evaluators may form incorrect beliefs about AI behavior in complex…

Artificial Intelligence · Computer Science 2025-10-22 Leon Lang , Patrick Forré

The ability to acquire latent semantics is one of the key properties that determines the performance of language models. One convenient approach to invoke this ability is to prepend metadata (e.g. URLs, domains, and styles) at the beginning…

As AI systems continue to evolve, their rigorous evaluation becomes crucial for their development and deployment. Researchers have constructed various large-scale benchmarks to determine their capabilities, typically against a gold-standard…

Computation and Language · Computer Science 2025-05-09 Yan Zhuang , Qi Liu , Zachary A. Pardos , Patrick C. Kyllonen , Jiyun Zu , Zhenya Huang , Shijin Wang , Enhong Chen

This survey paper chronicles the evolution of evaluation in multimodal artificial intelligence (AI), framing it as a progression of increasingly sophisticated "cognitive examinations." We argue that the field is undergoing a paradigm shift,…

Artificial Intelligence · Computer Science 2026-01-07 Mayank Ravishankara , Varindra V. Persad Maharaj

Recent published evidence from frontier laboratories shows that contemporary AI models can recognise evaluation contexts, latently represent them, and behave differently under those contexts than under deployment-continuous conditions.…

Artificial Intelligence · Computer Science 2026-05-13 Varad Vishwarupe , Nigel Shadbolt , Marina Jirotka , Ivan Flechais

Metacognition -- assessing the quality of one's own cognitive performance -- guides adaptive behavior across species. Substantial research demonstrates that confidence signals can be extracted from language model outputs, yet a fundamental…

Machine Learning · Computer Science 2026-05-20 Dharshan Kumaran , Nathaniel Daw , Simon Osindero , Petar Veličković , Viorica Patraucean

Emotions that somebody develops based on an argument do not only depend on the argument itself - they are also influenced by a subjective evaluation of the argument's potential impact on the self. For instance, an argument to ban plastic…

Computation and Language · Computer Science 2026-03-05 Lynn Greschner , Sabine Weber , Roman Klinger

Large Language Model (LLM) based judges form the underpinnings of key safety evaluation processes such as offline benchmarking, automated red-teaming, and online guardrailing. This widespread requirement raises the crucial question: can we…

Machine Learning · Computer Science 2025-03-07 Francisco Eiras , Eliott Zemour , Eric Lin , Vaikkunth Mugunthan

Prominent AI experts have suggested that companies developing high-risk AI systems should be required to show that such systems are safe before they can be developed or deployed. The goal of this paper is to expand on this idea and explore…

Computers and Society · Computer Science 2024-06-25 Akash R. Wasil , Joshua Clymer , David Krueger , Emily Dardaman , Simeon Campos , Evan R. Murphy

Despite remarkable achievements in artificial intelligence, the deployability of learning-enabled systems in high-stakes real-world environments still faces persistent challenges. For example, in safety-critical domains like autonomous…

Artificial Intelligence · Computer Science 2023-12-19 Minjae Cho , Chuangchuang Sun

Standard benchmarks of bias and fairness in large language models (LLMs) measure the association between the user attributes stated or implied by a prompt and the LLM's short text response, but human-AI interaction increasingly requires…

Computation and Language · Computer Science 2025-06-06 Kristian Lum , Jacy Reese Anthis , Kevin Robinson , Chirag Nagpal , Alexander D'Amour

Artificial intelligence systems are now deployed at scale across sectors, accompanied by a growing number of real-world incidents ranging from misinformation and cybercrime to autonomous-system failures. Databases of AI incidents index…

Computers and Society · Computer Science 2026-04-23 Sophia Abraham , Taiye Chen , Cyril Chhun , Giovanna Jaramillo-Gutierrez , Simon Mylius , Sayash Raaj , Peter Slattery , Sean McGregor

Understanding toxicity in user conversations is undoubtedly an important problem. Addressing "covert" or implicit cases of toxicity is particularly hard and requires context. Very few previous studies have analysed the influence of…

Computation and Language · Computer Science 2022-10-19 Atijit Anuchitanukul , Julia Ive , Lucia Specia

Machine learning models for speech emotion recognition (SER) can be trained for different tasks and are usually evaluated based on a few available datasets per task. Tasks could include arousal, valence, dominance, emotional categories, or…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-13 Anna Derington , Hagen Wierstorf , Ali Özkil , Florian Eyben , Felix Burkhardt , Björn W. Schuller