English
Related papers

Related papers: Rater Cohesion and Quality from a Vicarious Perspe…

200 papers

Applications designed for entertainment and other non-instrumental purposes are challenging to optimize because the relationships between system parameters and user experience can be unclear. Ideally, we would crowdsource these design…

Human-Computer Interaction · Computer Science 2022-04-26 Alan Medlar , Jing Li , Yang Liu , Dorota Glowacka

Summarization systems are ultimately evaluated by human annotators and raters. Usually, annotators and raters do not reflect the demographics of end users, but are recruited through student populations or crowdsourcing platforms with skewed…

Computation and Language · Computer Science 2021-10-12 Anna Jørgensen , Anders Søgaard

Machine learning approaches often require training and evaluation datasets with a clear separation between positive and negative examples. This risks simplifying and even obscuring the inherent subjectivity present in many tasks. Preserving…

Human-Computer Interaction · Computer Science 2023-06-21 Lora Aroyo , Alex S. Taylor , Mark Diaz , Christopher M. Homan , Alicia Parrish , Greg Serapio-Garcia , Vinodkumar Prabhakaran , Ding Wang

Collecting annotations from human raters often results in a trade-off between the quantity of labels one wishes to gather and the quality of these labels. As such, it is often only possible to gather a small amount of high-quality labels.…

Machine Learning · Computer Science 2021-10-05 Neel Nanda , Jonathan Uesato , Sven Gowal

Incorporating every annotator's perspective is crucial for unbiased data modeling. Annotator fatigue and changing opinions over time can distort dataset annotations. To combat this, we propose to learn a more accurate representation of…

Machine Learning · Computer Science 2024-06-05 Uthman Jinadu , Yi Ding

Research into community content moderation often assumes that moderation teams govern with a single, unified voice. However, recent work has found that moderators disagree with one another at modest, but concerning rates. The problem is not…

Human-Computer Interaction · Computer Science 2024-11-01 Vinay Koshy , Frederick Choi , Yi-Shyuan Chiang , Hari Sundaram , Eshwar Chandrasekharan , Karrie Karahalios

The voluntary process of Wikipedia edition provides an environment where the outcome is clearly a collective product of interactions involving a large number of people. We propose a simple agent-based model, developed from real data, to…

Physics and Society · Physics 2015-06-22 Y. Gandica , F. Sampaio dos Aidos , J. Carvalho

The rise of online platforms exacerbated the spread of hate speech, demanding scalable and effective detection. However, the accuracy of hate speech detection systems heavily relies on human-labeled data, which is inherently susceptible to…

Computation and Language · Computer Science 2025-06-13 Tommaso Giorgi , Lorenzo Cima , Tiziano Fagni , Marco Avvenuti , Stefano Cresci

Hate speech moderation remains a challenging task for social media platforms. Human-AI collaborative systems offer the potential to combine the strengths of humans' reliability and the scalability of machine learning to tackle this issue…

Human-Computer Interaction · Computer Science 2023-07-25 Philippe Lammerts , Philip Lippmann , Yen-Chia Hsu , Fabio Casati , Jie Yang

This study uses the cosine similarity ratio, embedding regression, and manual re-annotation to diagnose hate speech classification. We begin by computing cosine similarity ratio on a dataset "Measuring Hate Speech" that contains 135,556…

Computation and Language · Computer Science 2024-11-27 Xilin Yang

When annotators disagree, predicting the labels given by individual annotators can capture nuances overlooked by traditional label aggregation. We introduce three approaches to predicting individual annotator ratings on the toxicity of text…

Computation and Language · Computer Science 2024-10-17 Harbani Jaggi , Kashyap Murali , Eve Fleisig , Erdem Bıyık

Providing constructive feedback to paper authors is a core component of peer review. With reviewers increasingly having less time to perform reviews, automated support systems are required to ensure high reviewing quality, thus making the…

Computation and Language · Computer Science 2025-09-23 Abdelrahman Sadallah , Tim Baumgärtner , Iryna Gurevych , Ted Briscoe

Fairness in recommender systems has recently received attention from researchers. Unfair recommendations have negative impact on the effectiveness of recommender systems as it may degrade users' satisfaction, loyalty, and at worst, it can…

Information Retrieval · Computer Science 2019-11-05 Masoud Mansoury , Himan Abdollahpouri , Joris Rombouts , Mykola Pechenizkiy

Current multimodal toxicity benchmarks typically use a single binary hatefulness label. This coarse approach conflates two fundamentally different characteristics of expression: tone and content. Drawing on communication science theory, we…

Computation and Language · Computer Science 2026-03-25 Nils A. Herrmann , Tobias Eder , Jingyi He , Georg Groh

Human annotations are an important source of information in the development of natural language understanding approaches. As under the pressure of productivity annotators can assign different labels to a given text, the quality of produced…

Computation and Language · Computer Science 2020-10-29 Kristian Miok , Gregor Pirs , Marko Robnik-Sikonja

Recent studies comparing AI-generated and human-authored literary texts have produced conflicting results: some suggest AI already surpasses human quality, while others argue it still falls short. We start from the hypothesis that such…

Computation and Language · Computer Science 2025-06-05 Guillermo Marco , Julio Gonzalo , Víctor Fresno

The use of large language models like ChatGPT in code review offers promising efficiency gains but also raises concerns about correctness and safety. Existing evaluation methods for code review generation either rely on automatic…

Software Engineering · Computer Science 2025-12-18 Robert Heumüller , Frank Ortmeier

Human evaluation remains the primary standard for assessing modern AI systems, yet annotator disagreement, bias, and variability make system rankings fragile under standard majority vote aggregation. Majority vote discards annotator…

AI alignment relies on annotator judgments, yet annotation pipelines often treat annotators as interchangeable, obscuring how their social position shapes annotation. We introduce reflexive annotating as a probe that invites crowd workers…

Human-Computer Interaction · Computer Science 2026-04-22 Anne Arzberger , Celine Offerman , Ujwal Gadiraju , Alessandro Bozzon , Jie Yang

Community Notes is X's crowdsourced fact-checking program: contributors write short notes that add context to potentially misleading posts, and other contributors rate whether those notes are helpful. Its algorithm uses a matrix…

Social and Information Networks · Computer Science 2026-04-14 Mohak Goyal , Nishka Arora , Ashish Goel