中文
相关论文

相关论文: Learning from the Worst: Dynamically Generated Dat…

200 篇论文

Behavioural testing -- verifying system capabilities by validating human-designed input-output pairs -- is an alternative evaluation method of natural language processing systems proposed to address the shortcomings of the standard…

计算与语言 · 计算机科学 2022-07-05 Pedro Henrique Luz de Araujo , Benjamin Roth

Online propaganda poses a severe threat to the integrity of societies. However, existing datasets for detecting online propaganda have a key limitation: they were annotated using weak labels that can be noisy and even incorrect. To address…

计算与语言 · 计算机科学 2024-11-27 Abdurahman Maarouf , Dominik Bär , Dominique Geissler , Stefan Feuerriegel

Social media platforms and online streaming services have spawned a new breed of Hate Speech (HS). Due to the massive amount of user-generated content on these sites, modern machine learning techniques are found to be feasible and…

Social media, particularly Twitter, has seen a significant increase in incidents like trolling and hate speech. Thus, identifying hate speech is the need of the hour. This paper introduces a computational framework to curb the hate content…

计算与语言 · 计算机科学 2024-09-10 Anusha Chhabra , Dinesh Kumar Vishwakarma

Hate Speech takes many forms to target communities with derogatory comments, and takes humanity a step back in societal progress. HateXplain is a recently published and first dataset to use annotated spans in the form of rationales, along…

计算与语言 · 计算机科学 2022-08-10 Arvind Subramaniam , Aryan Mehra , Sayani Kundu

More capable language models increasingly saturate existing task benchmarks, in some cases outperforming humans. This has left little headroom with which to measure further progress. Adversarial dataset creation has been proposed as a…

计算与语言 · 计算机科学 2021-11-17 Jason Phang , Angelica Chen , William Huang , Samuel R. Bowman

Detecting hate speech in the workplace is a unique classification task, as the underlying social context implies a subtler version of conventional hate speech. Applications regarding a state-of the-art workplace sexism detection model…

计算与语言 · 计算机科学 2020-07-09 Dylan Grosz , Patricia Conde-Cespedes

Crowdsourcing has been the prevalent paradigm for creating natural language understanding datasets in recent years. A common crowdsourcing practice is to recruit a small number of high-quality workers, and have them massively generate…

计算与语言 · 计算机科学 2019-08-29 Mor Geva , Yoav Goldberg , Jonathan Berant

With the spread of social networks and their unfortunate use for hate speech, automatic detection of the latter has become a pressing problem. In this paper, we reproduce seven state-of-the-art hate speech detection models from prior work,…

计算与语言 · 计算机科学 2018-11-06 Tommi Gröndahl , Luca Pajola , Mika Juuti , Mauro Conti , N. Asokan

The existing research has primarily focused on text and image-based hate speech detection, video-based approaches remain underexplored. In this work, we introduce a novel dataset, ImpliHateVid, specifically curated for implicit hate speech…

计算机视觉与模式识别 · 计算机科学 2025-08-18 Mohammad Zia Ur Rehman , Anukriti Bhatnagar , Omkar Kabde , Shubhi Bansal , Nagendra Kumar

Hate speech has become pervasive in today's digital age. Although there has been considerable research to detect hate speech or generate counter speech to combat hateful views, these approaches still cannot completely eliminate the…

计算与语言 · 计算机科学 2023-10-24 Vibhor Agarwal , Yu Chen , Nishanth Sastry

The rise of online platforms exacerbated the spread of hate speech, demanding scalable and effective detection. However, the accuracy of hate speech detection systems heavily relies on human-labeled data, which is inherently susceptible to…

计算与语言 · 计算机科学 2025-06-13 Tommaso Giorgi , Lorenzo Cima , Tiziano Fagni , Marco Avvenuti , Stefano Cresci

Manually annotated datasets are crucial for training and evaluating Natural Language Processing models. However, recent work has discovered that even widely-used benchmark datasets contain a substantial number of erroneous annotations. This…

计算与语言 · 计算机科学 2023-06-01 Leon Weber , Barbara Plank

'Scale the model, scale the data, scale the compute' is the reigning sentiment in the world of generative AI today. While the impact of model scaling has been extensively studied, we are only beginning to scratch the surface of data scaling…

计算机与社会 · 计算机科学 2023-11-08 Abeba Birhane , Vinay Prabhu , Sang Han , Vishnu Naresh Boddeti , Alexandra Sasha Luccioni

Hate speech classification has been a long-standing problem in natural language processing. However, even though there are numerous hate speech detection methods, they usually overlook a lot of hateful statements due to them being implicit…

计算与语言 · 计算机科学 2022-08-30 Debaditya Pal , Kaustubh Chaudhari , Harsh Sharma

Accurate detection and classification of online hate is a difficult task. Implicit hate is particularly challenging as such content tends to have unusual syntax, polysemic words, and fewer markers of prejudice (e.g., slurs). This problem is…

计算与语言 · 计算机科学 2021-06-11 Austin Botelho , Bertie Vidgen , Scott A. Hale

Hate speech is harmful content that directly attacks or promotes hatred against members of groups or individuals based on actual or perceived aspects of identity, such as racism, religion, or sexual orientation. This can affect social life…

计算与语言 · 计算机科学 2024-03-19 Arijit Das , Somashree Nandy , Rupam Saha , Srijan Das , Diganta Saha

We conduct relatively extensive investigations of automatic hate speech (HS) detection using different state-of-the-art (SoTA) baselines over 11 subtasks of 6 different datasets. Our motivation is to determine which of the recent SoTA…

计算与语言 · 计算机科学 2022-10-12 Tosin Adewumi , Sana Sabah Sabry , Nosheen Abid , Foteini Liwicki , Marcus Liwicki

We introduce a new large-scale NLI benchmark dataset, collected via an iterative, adversarial human-and-model-in-the-loop procedure. We show that training models on this new dataset leads to state-of-the-art performance on a variety of…

计算与语言 · 计算机科学 2020-05-07 Yixin Nie , Adina Williams , Emily Dinan , Mohit Bansal , Jason Weston , Douwe Kiela

We present our experience as annotators in the creation of high-quality, adversarial machine-reading-comprehension data for extractive QA for Task 1 of the First Workshop on Dynamic Adversarial Data Collection (DADC). DADC is an emergent…

计算与语言 · 计算机科学 2022-06-30 Damian Y. Romero Diaz , Magdalena Anioł , John Culnan