English
Related papers

Related papers: Rater Cohesion and Quality from a Vicarious Perspe…

200 papers

Sentiment analysis is often a crowdsourcing task prone to subjective labels given by many annotators. It is not yet fully understood how the annotation bias of each annotator can be modeled correctly with state-of-the-art methods. However,…

This paper investigates how collaborative AI systems can enhance user agency in identifying and evaluating misinformation on social media platforms. Traditional methods, such as personal judgment or basic fact-checking, often fall short…

Human-Computer Interaction · Computer Science 2025-07-01 Varun Sangwan , Heidi Makitalo

Expertise of annotators has a major role in crowdsourcing based opinion aggregation models. In such frameworks, accuracy and biasness of annotators are occasionally taken as important features and based on them priority of the annotators…

Human-Computer Interaction · Computer Science 2017-09-01 Sujoy Chatterjee , Anirban Mukhopadhyay , Malay Bhattacharyya

Prior research in psychology has found that people's decisions are often inconsistent. An individual's decisions vary across time, and decisions vary even more across people. Inconsistencies have been identified not only in subjective…

Human-Computer Interaction · Computer Science 2024-07-17 Nina Grgić-Hlača , Junaid Ali , Krishna P. Gummadi , Jennifer Wortman Vaughan

Autoraters, also referred to as LLM-as-judges, are increasingly used for evaluation and automated content moderation. However, there is limited statistical analysis of how modifications in a rubric presented to both humans and autoraters…

Computation and Language · Computer Science 2026-05-08 Jessica Huynh , Alfredo Gomez , Athiya Deviyani , Renee Shelby , Jeffrey P. Bigham , Fernando Diaz

Speech enhancement techniques improve the quality or the intelligibility of an audio signal by removing unwanted noise. It is used as preprocessing in numerous applications such as speech recognition, hearing aids, broadcasting and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-05 Angélica S. Z. Suárez , Clément Laroche , Line H. Clemmensen , Sneha Das

A major challenge in the field of Text Generation is evaluation: Human evaluations are cost-intensive, and automated metrics often display considerable disagreement with human judgments. In this paper, we propose a statistical model of Text…

Computation and Language · Computer Science 2023-06-07 Jan Deriu , Pius von Däniken , Don Tuggener , Mark Cieliebak

This work is motivated by the need to assess the degree of agreement between two independent groups of raters. It proposes two new methods.

Applications · Statistics 2018-06-18 Madhusmita Panda , Sharayu Paranjpe , Anil Gore

As a major source for information on virtually any topic, Wikipedia serves an important role in public dissemination and consumption of knowledge. As a result, it presents tremendous potential for people to promulgate their own points of…

Social and Information Networks · Computer Science 2011-11-10 Sanmay Das , Allen Lavoie , Malik Magdon-Ismail

Human behaviors are often guided or constrained by social norms, which are defined as shared, commonsense rules. For example, underlying an action ``\textit{report a witnessed crime}" are social norms that inform our conduct, such as…

Computers and Society · Computer Science 2025-12-19 Yuxi Sun , Wei Gao , Hongzhan Lin , Jing Ma , Wenxuan Zhang

Automatic dialogue coherence evaluation has attracted increasing attention and is crucial for developing promising dialogue systems. However, existing metrics have two major limitations: (a) they are mostly trained in a simplified two-level…

Computation and Language · Computer Science 2021-07-23 Zheng Ye , Liucun Lu , Lishan Huang , Liang Lin , Xiaodan Liang

Human feedback has become the de facto standard for evaluating the performance of Large Language Models, and is increasingly being used as a training objective. However, it is not clear which properties of a generated output this single…

Computation and Language · Computer Science 2024-01-17 Tom Hosking , Phil Blunsom , Max Bartolo

In this paper, we explore the feasibility of leveraging large language models (LLMs) to automate or otherwise assist human raters with identifying harmful content including hate speech, harassment, violent extremism, and election…

Preference elicitation frameworks feature heavily in the research on participatory ethical AI tools and provide a viable mechanism to enquire and incorporate the moral values of various stakeholders. As part of the elicitation process,…

Computers and Society · Computer Science 2024-08-07 Kyle Boerstler , Vijay Keswani , Lok Chan , Jana Schaich Borg , Vincent Conitzer , Hoda Heidari , Walter Sinnott-Armstrong

How do Large Language Models understand moral dimensions compared to humans? This first large-scale Bayesian evaluation of market-leading language models provides the answer. In contrast to prior work using deterministic ground truth…

Computation and Language · Computer Science 2025-11-24 Maciej Skorski , Alina Landowska

Qualitative analysis is critical to understanding human datasets in many social science disciplines. A central method in this process is inductive coding, where researchers identify and interpret codes directly from the datasets themselves.…

Computation and Language · Computer Science 2026-04-21 John Chen , Alexandros Lotsos , Sihan Cheng , Caiyi Wang , Lexie Zhao , Yanjia Zhang , Jessica Hullman , Bruce Sherin , Uri Wilensky , Michael Horn

Community-based fact-checking systems, such as Community Notes on X (formerly Twitter), aim to mitigate online misinformation by surfacing annotations judged helpful by contributors with diverse viewpoints. While prior work has shown that…

Social and Information Networks · Computer Science 2026-01-21 Yuwei Chuai , Gabriele Lenzini , Nicolas Pröllochs

Even though considerable attention has been given to the polarity of words (positive and negative) and the creation of large polarity lexicons, research in emotion analysis has had to rely on limited and small emotion lexicons. In this…

Computation and Language · Computer Science 2013-08-30 Saif M. Mohammad , Peter D. Turney

The detection and identification of toxic comments are conducive to creating a civilized and harmonious Internet environment. In this experiment, we collected various data sets related to toxic comments. Because of the characteristics of…

Computation and Language · Computer Science 2022-03-08 Zhichang Wang , Qipeng Zhu

Content annotation at scale remains challenging, requiring substantial human expertise and effort. This paper presents a case study in code documentation analysis, where we explore the balance between automation efficiency and annotation…

Human-Computer Interaction · Computer Science 2025-04-29 Mingyue Yuan , Jieshan Chen , Zhenchang Xing , Gelareh Mohammadi , Aaron Quigley