English
Related papers

Related papers: How We Define Harm Impacts Data Annotations: Expla…

200 papers

Safety policies define what constitutes safe and unsafe AI outputs, guiding data annotation and model development. However, annotation disagreement is pervasive and can stem from multiple sources such as operational failures (annotators…

Artificial Intelligence · Computer Science 2026-05-08 Alex Oesterling , Donghao Ren , Yannick Assogba , Dominik Moritz , Sunnie S. Y. Kim , Leon Gatys , Fred Hohman

Hate speech is commonly defined as any communication that disparages a target group of people based on some characteristic such as race, colour, ethnicity, gender, sexual orientation, nationality, religion, or other characteristic. Due to…

Computation and Language · Computer Science 2018-09-13 Ona de Gibert , Naiara Perez , Aitor García-Pablos , Montse Cuadros

Adult content detection still poses a great challenge for automation. Existing classifiers primarily focus on distinguishing between erotic and non-erotic texts. However, they often need more nuance in assessing the potential harm.…

Computation and Language · Computer Science 2023-10-24 Inez Okulska , Emilia Wiśnios

Collecting annotations from human raters often results in a trade-off between the quantity of labels one wishes to gather and the quality of these labels. As such, it is often only possible to gather a small amount of high-quality labels.…

Machine Learning · Computer Science 2021-10-05 Neel Nanda , Jonathan Uesato , Sven Gowal

Hate speech is a widespread and harmful form of online discourse, encompassing slurs and defamatory posts that can have serious social, psychological, and sometimes physical impacts on targeted individuals and communities. As social media…

Machine Learning · Computer Science 2025-08-08 Santosh Chapagain , Shah Muhammad Hamdi , Soukaina Filali Boubrahimi

Now-a-days, derogatory comments are often made by one another, not only in offline environment but also immensely in online environments like social networking websites and online communities. So, an Identification combined with Prevention…

Computation and Language · Computer Science 2019-03-19 Navoneel Chakrabarty

Hate speech, offensive language, aggression, racism, sexism, and other abusive language are common phenomena in social media. There is a need for Artificial Intelligence(AI)based intervention which can filter hate content at scale. Most…

Computation and Language · Computer Science 2024-11-13 Prashant Kapil , Asif Ekbal

The sheer volume of online user-generated content has rendered content moderation technologies essential in order to protect digital platform audiences from content that may cause anxiety, worry, or concern. Despite the efforts towards…

Computer Vision and Pattern Recognition · Computer Science 2022-12-02 Ioannis Sarridis , Christos Koutlis , Olga Papadopoulou , Symeon Papadopoulos

The proliferation of harmful content on online social media platforms has necessitated empirical understandings of experiences of harm online and the development of practices for harm mitigation. Both understandings of harm and approaches…

Human-Computer Interaction · Computer Science 2021-09-20 Morgan Klaus Scheuerman , Jialun Aaron Jiang , Casey Fiesler , Jed R. Brubaker

Recent research at the intersection of AI explainability and fairness has focused on how explanations can improve human-plus-AI task performance as assessed by fairness measures. We propose to characterize what constitutes an explanation…

Computation and Language · Computer Science 2023-10-24 Tin Nguyen , Jiannan Xu , Aayushi Roy , Hal Daumé , Marine Carpuat

When humans judge the affective content of texts, they also implicitly assess the correctness of such judgment, that is, their confidence. We hypothesize that people's (in)confidence that they performed well in an annotation task leads to…

Computation and Language · Computer Science 2021-03-03 Enrica Troiano , Sebastian Padó , Roman Klinger

Dark humor often relies on subtle cultural nuances and implicit cues that require contextual reasoning to interpret, posing safety challenges that current static benchmarks fail to capture. To address this, we introduce a novel multimodal,…

Computation and Language · Computer Science 2026-03-20 Ahmed Sharshar , Hosam Elgendy , Saad El Dine Ahmed , Yasser Rohaim , Yuxia Wang

Organizations worldwide that rely on data-driven approaches regularly employ forecasting methods to enhance their planning and decision-making processes. While extensive research has examined the harms associated with traditional machine…

Other Statistics · Statistics 2025-03-14 Bahman Rostami-Tabar , Travis Greene , Galit Shmueli , Rob J. Hyndman

The proliferation of online hate speech has necessitated the creation of algorithms which can detect toxicity. Most of the past research focuses on this detection as a classification task, but assigning an absolute toxicity label is often…

Computation and Language · Computer Science 2022-06-28 Millon Madhur Das , Punyajoy Saha , Mithun Das

Large language models (LLMs) increasingly operate on long inputs, yet their behavior when harmful sentences are sparsely embedded within such inputs remains poorly understood. We present a sensitivity analysis that probes how LLMs extract…

Computation and Language · Computer Science 2026-05-27 Faeze Ghorbanpour , Alexander Fraser

The age of social media is flooded with Internet memes, necessitating a clear grasp and effective identification of harmful ones. This task presents a significant challenge due to the implicit meaning embedded in memes, which is not…

Computation and Language · Computer Science 2024-01-25 Hongzhan Lin , Ziyang Luo , Wei Gao , Jing Ma , Bo Wang , Ruichao Yang

Understanding toxicity in user conversations is undoubtedly an important problem. Addressing "covert" or implicit cases of toxicity is particularly hard and requires context. Very few previous studies have analysed the influence of…

Computation and Language · Computer Science 2022-10-19 Atijit Anuchitanukul , Julia Ive , Lucia Specia

The dissemination of online hate speech can have serious negative consequences for individuals, online communities, and entire societies. This and the large volume of hateful online content prompted both practitioners', i.e., in content…

Computation and Language · Computer Science 2025-04-14 Julian Bäumler , Louis Blöcher , Lars-Joel Frey , Xian Chen , Markus Bayer , Christian Reuter

In this work, we explore the capability of Large Language Models (LLMs) to annotate hate speech and abusiveness while considering predefined annotator personas within the strong-to-weak data perspectivism spectra. We evaluated LLM-generated…

Computation and Language · Computer Science 2025-08-26 Olufunke O. Sarumi , Charles Welch , Daniel Braun , Jörg Schlötterer

Supervised approaches generally rely on majority-based labels. However, it is hard to achieve high agreement among annotators in subjective tasks such as hate speech detection. Existing neural network models principally regard labels as…

Computation and Language · Computer Science 2023-01-11 Wenjie Yin , Vibhor Agarwal , Aiqi Jiang , Arkaitz Zubiaga , Nishanth Sastry
‹ Prev 1 4 5 6 7 8 10 Next ›