English
Related papers

Related papers: Beyond Agreement: Rethinking Ground Truth in Educa…

200 papers

Generative Artificial Intelligence (GenAI) is now widespread in education, yet the efficacy of GenAI systems remains constrained by the quality and interpretation of the labeled data used to train and evaluate them. Studies commonly report…

Computers and Society · Computer Science 2026-04-01 Danielle R. Thomas , Conrad Borchers , Kirk P. Vanacore , Kenneth R. Koedinger , René F. Kizilcec

Many machine learning systems today are trained on large amounts of human-annotated data. Data annotation tasks that require a high level of competency make data acquisition expensive, while the resulting labels are often subjective,…

Machine Learning · Computer Science 2020-04-08 Emmanouil Antonios Platanios , Maruan Al-Shedivat , Eric Xing , Tom Mitchell

In machine learning, "ground truth" refers to the assumed correct labels used to train and evaluate models. However, the foundational "ground truth" paradigm rests on a positivistic fallacy that treats human disagreement as technical noise…

Artificial Intelligence · Computer Science 2026-04-28 Sheza Munir , Benjamin Mah , Krisha Kalsi , Shivani Kapania , Julian Posada , Edith Law , Ding Wang , Syed Ishtiaque Ahmed

Labelled "ground truth" datasets are routinely used to evaluate and audit AI algorithms applied in high-stakes settings. However, there do not exist widely accepted benchmarks for the quality of labels in these datasets. We provide…

Computation and Language · Computer Science 2021-11-18 Abhilash Mishra , Yash Gorana

We present an empirical study of how both experienced tutors and non-tutors judge the correctness of tutor praise responses under different Artificial Intelligence (AI)-assisted interfaces, types of explanation (textual explanations vs.…

Human-Computer Interaction · Computer Science 2026-01-06 Eason Chen , Jeffrey Li , Scarlett Huang , Xinyi Tang , Jionghao Lin , Paulo Carvalho , Kenneth Koedinger

In supervised learning, low quality annotations lead to poorly performing classification and detection models, while also rendering evaluation unreliable. This is particularly apparent on temporal data, where annotation quality is affected…

Artificial intelligence is applied in a range of sectors, and is relied upon for decisions requiring a high level of trust. For regression methods, trust is increased if they approximate the true input-output relationships and perform…

Machine Learning · Computer Science 2021-05-05 A. I. Parkes , A. J. Sobey , D. A. Hudson

As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworthiness of these systems is paramount. However, AI is currently facing a reproducibility crisis driven…

Machine Learning · Computer Science 2026-05-14 Deepak Pandita , Flip Korn , Chris Welty , Christopher M. Homan

Human annotation remains the foundation of reliable and interpretable data in Natural Language Processing (NLP). As annotation and evaluation tasks continue to expand, from categorical labelling to segmentation, subjective judgment, and…

Computation and Language · Computer Science 2026-04-02 Joseph James

"Gold" and "ground truth" human-mediated labels have error. The effects of this error can escape commonly reported metrics of label quality or obscure questions of accuracy, bias, fairness, and usefulness during model evaluation. This study…

Computation and Language · Computer Science 2024-11-26 Michael Hardy

The use of machine learning (ML)-based language models (LMs) to monitor content online is on the rise. For toxic text identification, task-specific fine-tuning of these models are performed using datasets labeled by annotators who provide…

Computation and Language · Computer Science 2021-12-08 Kofi Arhin , Ioana Baldini , Dennis Wei , Karthikeyan Natesan Ramamurthy , Moninder Singh

Explanation methods in Interpretable NLP often explain the model's decision by extracting evidence (rationale) from the input texts supporting the decision. Benchmark datasets for rationales have been released to evaluate how good the…

Computation and Language · Computer Science 2022-04-12 Cheng-Han Chiang , Hung-yi Lee

Human variation in labeling is often considered noise. Annotation projects for machine learning (ML) aim at minimizing human label variation, with the assumption to maximize data quality and in turn optimize and maximize machine learning…

Computation and Language · Computer Science 2022-11-07 Barbara Plank

Crowdsourcing has become a common approach for annotating large amounts of data. It has the advantage of harnessing a large workforce to produce large amounts of data in a short time, but comes with the disadvantage of employing non-expert…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-12 Irene Martin-Morato , Annamaria Mesaros

Current supervised deep learning frameworks rely on annotated data for modeling the underlying data distribution of a given task. In particular for computer vision algorithms powered by deep learning, the quality of annotated data is the…

Computer Vision and Pattern Recognition · Computer Science 2019-12-24 Joseph Nassar , Viveca Pavon-Harr , Marc Bosch , Ian McCulloh

The annotation of domain experts is important for some medical applications where the objective ground truth is ambiguous to define, e.g., the rehabilitation for some chronic diseases, and the prescreening of some musculoskeletal…

Machine Learning · Computer Science 2023-03-10 Chongyang Wang , Yuan Gao , Chenyou Fan , Junjie Hu , Tin Lun Lam , Nicholas D. Lane , Nadia Bianchi-Berthouze

Pairwise preferences over model responses are widely collected to evaluate and provide feedback to large language models (LLMs). Given two alternative model responses to the same input, a human or AI annotator selects the "better" response.…

Computation and Language · Computer Science 2025-07-24 Arduin Findeis , Floris Weers , Guoli Yin , Ke Ye , Ruoming Pang , Tom Gunter

Supervised machine learning assumes that labeled data provide accurate measurements of the concepts models are meant to learn. Yet in practice, human labeling introduces systematic variation arising from ambiguous items, divergent…

Methodology · Statistics 2026-04-10 Robert Chew , Stephanie Eckman , Christoph Kern , Frauke Kreuter

In many settings, an effective way of evaluating objects of interest is to collect evaluations from dispersed individuals and to aggregate these evaluations together. Some examples are categorizing online content and evaluating student…

Computer Science and Game Theory · Computer Science 2016-06-23 Alice Gao , James R. Wright , Kevin Leyton-Brown

The process of gathering ground truth data through human annotation is a major bottleneck in the use of information extraction methods for populating the Semantic Web. Crowdsourcing-based approaches are gaining popularity in the attempt to…

Human-Computer Interaction · Computer Science 2022-09-21 Anca Dumitrache , Oana Inel , Benjamin Timmermans , Carlos Ortiz , Robert-Jan Sips , Lora Aroyo , Chris Welty
‹ Prev 1 2 3 10 Next ›