English
Related papers

Related papers: Evaluating and Characterizing Human Rationales

200 papers

Traditionally, the way one evaluates the performance of an Artificial Intelligence (AI) system is via a comparison to human performance in specific tasks, treating humans as a reference for high-level cognition. However, these comparisons…

Artificial Intelligence · Computer Science 2019-11-25 Camilo M. Signorelli , Xerxes D. Arsiwalla

At the heart of improving conversational AI is the open problem of how to evaluate conversations. Issues with automatic metrics are well known (Liu et al., 2016, arXiv:1603.08023), with human evaluations still considered the gold standard.…

Computation and Language · Computer Science 2022-01-14 Eric Michael Smith , Orion Hsu , Rebecca Qian , Stephen Roller , Y-Lan Boureau , Jason Weston

People are known to judge artificial intelligence using a utilitarian moral philosophy and humans using a moral philosophy emphasizing perceived intentions. But why do people judge humans and machines differently? Psychology suggests that…

Computers and Society · Computer Science 2023-09-20 Jingling Zhang , Jane Conway , César A. Hidalgo

Large language models have recently shown promising progress in mathematical reasoning when fine-tuned with human-generated sequences walking through a sequence of solution steps. However, the solution sequences are not formally structured…

Machine Learning · Computer Science 2022-12-07 Andrew J. Nam , Mengye Ren , Chelsea Finn , James L. McClelland

Mathematical reasoning---a core ability within human intelligence---presents some unique challenges as a domain: we do not come to understand and solve mathematical problems primarily on the back of experience and evidence, but on the basis…

Machine Learning · Computer Science 2019-04-03 David Saxton , Edward Grefenstette , Felix Hill , Pushmeet Kohli

Value learning is a crucial aspect of safe and ethical AI. This is primarily pursued by methods inferring human values from behaviour. However, humans care about much more than we are able to demonstrate through our actions. Consequently,…

Artificial Intelligence · Computer Science 2025-05-28 Paul de Font-Reaulx

Due to flourish of the Web 2.0, web opinion sources are rapidly emerging containing precious information useful for both customers and manufactures. Recently, feature based opinion mining techniques are gaining momentum in which customer…

Information Retrieval · Computer Science 2015-04-27 Ahmad Kamal

If artificial intelligence (AI) is to be applied in safety-critical domains, its performance needs to be evaluated reliably. The present study aimed to understand how humans evaluate AI systems for person detection in automatic train…

Human-Computer Interaction · Computer Science 2025-04-04 Romy Müller

Automated decision systems increasingly rely on human oversight to ensure accuracy in uncertain cases. This paper presents a practical framework for optimizing such human-in-the-loop classification systems using a double-threshold policy.…

Human-Computer Interaction · Computer Science 2026-01-13 Goran Muric , Steven Minton

While automatic performance metrics are crucial for machine learning of artificial human-like behaviour, the gold standard for evaluation remains human judgement. The subjective evaluation of artificial human-like behaviour in embodied…

Human-Computer Interaction · Computer Science 2021-08-16 Pieter Wolfert , Jeffrey M. Girard , Taras Kucherenko , Tony Belpaeme

If machine learning models were to achieve superhuman abilities at various reasoning or decision-making tasks, how would we go about evaluating such models, given that humans would necessarily be poor proxies for ground truth? In this…

Machine Learning · Computer Science 2023-10-20 Lukas Fluri , Daniel Paleka , Florian Tramèr

Human evaluation is indispensable and inevitable for assessing the quality of texts generated by machine learning models or written by humans. However, human evaluation is very difficult to reproduce and its quality is notoriously unstable,…

Computation and Language · Computer Science 2023-05-04 Cheng-Han Chiang , Hung-yi Lee

Evaluating automatically generated text is generally hard due to the inherently subjective nature of many aspects of the output quality. This difficulty is compounded in automatic consultation note generation by differing opinions between…

Computation and Language · Computer Science 2022-11-18 Aleksandar Savkov , Francesco Moramarco , Alex Papadopoulos Korfiatis , Mark Perera , Anya Belz , Ehud Reiter

Many data mining approaches aim at modelling and predicting human behaviour. An important quantity of interest is the quality of model-based predictions, e.g. for finding a competition winner with best prediction performance. In real life,…

Human-Computer Interaction · Computer Science 2017-02-27 Kevin Jasberg , Sergej Sizov

Although text style transfer has witnessed rapid development in recent years, there is as yet no established standard for evaluation, which is performed using several automatic metrics, lacking the possibility of always resorting to human…

Computation and Language · Computer Science 2022-04-18 Huiyuan Lai , Jiali Mao , Antonio Toral , Malvina Nissim

One of the most crucial issues in data mining is to model human behaviour in order to provide personalisation, adaptation and recommendation. This usually involves implicit or explicit knowledge, either by observing user interactions, or by…

Human-Computer Interaction · Computer Science 2017-08-21 Kevin Jasberg , Sergej Sizov

Large language models (LMs) are capable of generating free-text rationales to aid question answering. However, prior work 1) suggests that useful self-rationalization is emergent only at significant scales (e.g., 175B parameter GPT-3); and…

Computation and Language · Computer Science 2024-05-24 Sahana Ramnath , Brihi Joshi , Skyler Hallinan , Ximing Lu , Liunian Harold Li , Aaron Chan , Jack Hessel , Yejin Choi , Xiang Ren

Automatic evaluation of language generation systems is a well-studied problem in Natural Language Processing. While novel metrics are proposed every year, a few popular metrics remain as the de facto metrics to evaluate tasks such as image…

Computation and Language · Computer Science 2020-10-27 Ozan Caglayan , Pranava Madhyastha , Lucia Specia

In this paper, we argue that the prevailing approach to training and evaluating machine learning models often fails to consider their real-world application within organizational or societal contexts, where they are intended to create…

Machine Learning · Computer Science 2025-04-24 Burcu Sayin , Jie Yang , Xinyue Chen , Andrea Passerini , Fabio Casati

Large language models (LLMs) are proficient at generating fluent text with minimal task-specific supervision. Yet, their ability to provide well-grounded rationalizations for knowledge-intensive tasks remains under-explored. Such tasks,…

Computation and Language · Computer Science 2024-02-02 Aditi Mishra , Sajjadur Rahman , Hannah Kim , Kushan Mitra , Estevam Hruschka