English
Related papers

Related papers: Annotation Sensitivity: Training Data Collection M…

200 papers

Laughter is considered one of the most overt signals of joy. Laughter is well-recognized as a multimodal phenomenon but is most commonly detected by sensing the sound of laughter. It is unclear how perception and annotation of laughter…

Sound · Computer Science 2022-11-03 Jose Vargas-Quiros , Laura Cabrera-Quiros , Catharine Oertel , Hayley Hung

Data annotated by humans is a source of knowledge by describing the peculiarities of the problem and therefore fueling the decision process of the trained model. Unfortunately, the annotation process for subjective natural language…

Computation and Language · Computer Science 2023-12-14 Kamil Kanclerz , Julita Bielaniewicz , Marcin Gruza , Jan Kocon , Stanisław Woźniak , Przemysław Kazienko

Hate speech detection models are only as good as the data they are trained on. Datasets sourced from social media suffer from systematic gaps and biases, leading to unreliable models with simplistic decision boundaries. Adversarial…

Computation and Language · Computer Science 2024-03-29 Janis Goldzycher , Paul Röttger , Gerold Schneider

It is increasingly recognized that human annotators do not always agree, and such disagreement is inherent in many annotation tasks. However, not all instances in a given task elicit the same degree of opinion divergence. In this paper, we…

Computation and Language · Computer Science 2026-05-05 Leixin Zhang , Çağrı Çöltekin

Well-annotated data is a prerequisite for good Natural Language Processing models. Too often, though, annotation decisions are governed by optimizing time or annotator agreement. We make a case for nuanced efforts in an interdisciplinary…

Computation and Language · Computer Science 2022-10-31 Federico Bianchi , Stefanie Anja Hills , Patricia Rossini , Dirk Hovy , Rebekah Tromble , Nava Tintarev

Natural language processing research has begun to embrace the notion of annotator subjectivity, motivated by variations in labelling. This approach understands each annotator's view as valid, which can be highly suitable for tasks that…

Computation and Language · Computer Science 2024-03-05 Amanda Cercas Curry , Gavin Abercrombie , Zeerak Talat

The development of real-time affect detection models often depends upon obtaining annotated data for supervised learning by employing human experts to label the student data. One open question in annotating affective data for affect…

Human-Computer Interaction · Computer Science 2019-01-15 Eda Okur , Sinem Aslan , Nese Alyuz , Asli Arslan Esme , Ryan S. Baker

Supervised classification heavily depends on datasets annotated by humans. However, in subjective tasks such as toxicity classification, these annotations often exhibit low agreement among raters. Annotations have commonly been aggregated…

Computation and Language · Computer Science 2024-05-17 Negar Mokhberian , Myrl G. Marmarelis , Frederic R. Hopp , Valerio Basile , Fred Morstatter , Kristina Lerman

The NLP community has long advocated for the construction of multi-annotator datasets to better capture the nuances of language interpretation, subjectivity, and ambiguity. This paper conducts a retrospective study to show how performance…

Computation and Language · Computer Science 2023-10-24 Pritam Kadasi , Mayank Singh

High-quality human annotations are necessary to create effective machine learning systems for social media. Low-quality human annotations indirectly contribute to the creation of inaccurate or biased learning systems. We show that human…

Social and Information Networks · Computer Science 2019-07-18 Rahul Pandey , Carlos Castillo , Hemant Purohit

Social media platforms provide users the freedom of expression and a medium to exchange information and express diverse opinions. Unfortunately, this has also resulted in the growth of abusive content with the purpose of discriminating…

Computation and Language · Computer Science 2021-07-01 Sohail Akhtar , Valerio Basile , Viviana Patti

Generated hateful and toxic content by a portion of users in social media is a rising phenomenon that motivated researchers to dedicate substantial efforts to the challenging direction of hateful content identification. We not only need an…

Social and Information Networks · Computer Science 2019-10-29 Marzieh Mozafari , Reza Farahbakhsh , Noel Crespi

Building a benchmark dataset for hate speech detection presents various challenges. Firstly, because hate speech is relatively rare, random sampling of tweets to annotate is very inefficient in finding hate speech. To address this, prior…

Computation and Language · Computer Science 2021-11-11 Md Mustafizur Rahman , Dinesh Balakrishnan , Dhiraj Murthy , Mucahid Kutlu , Matthew Lease

Online social media is rife with offensive and hateful comments, prompting the need for their automatic detection given the sheer amount of posts created every second. Creating high-quality human-labelled datasets for this task is difficult…

Computation and Language · Computer Science 2023-08-01 João A. Leite , Carolina Scarton , Diego F. Silva

Online abusive behavior is an important issue that breaks the cohesiveness of online social communities and even raises public safety concerns in our societies. Motivated by this rising issue, researchers have proposed, collected, and…

Social and Information Networks · Computer Science 2020-06-25 Md Rabiul Awal , Rui Cao , Roy Ka-Wei Lee , Sandra Mitrović

Content moderation and toxicity classification represent critical tasks with significant social implications. However, studies have shown that major classification models exhibit tendencies to magnify or reduce biases and potentially…

Computation and Language · Computer Science 2024-11-28 Haniyeh Ehsani Oskouie , Christina Chance , Claire Huang , Margaret Capetz , Elizabeth Eyeson , Majid Sarrafzadeh

Manual annotations are a prerequisite for many applications of machine learning. However, weaknesses in the annotation process itself are easy to overlook. In particular, scholars often choose what information to give to annotators without…

Social and Information Networks · Computer Science 2017-08-22 Kenneth Joseph , Lisa Friedland , William Hobbs , Oren Tsur , David Lazer

Hand-annotated data can vary due to factors such as subjective differences, intra-rater variability, and differing annotator expertise. We study annotations from different experts who labelled the same behavior classes on a set of animal…

Machine Learning · Computer Science 2021-06-14 Megan Tjandrasuwita , Jennifer J. Sun , Ann Kennedy , Swarat Chaudhuri , Yisong Yue

Some users of social media are spreading racist, sexist, and otherwise hateful content. For the purpose of training a hate speech detection system, the reliability of the annotations is crucial, but there is no universally agreed-upon…

Computation and Language · Computer Science 2017-01-30 Björn Ross , Michael Rist , Guillermo Carbonell , Benjamin Cabrera , Nils Kurowsky , Michael Wojatzki

Subjectivity and difference of opinion are key social phenomena, and it is crucial to take these into account in the annotation and detection process of derogatory textual content. In this paper, we use four datasets provided by…

Computation and Language · Computer Science 2023-05-03 Sadat Shahriar , Thamar Solorio