中文
相关论文

相关论文: SoftHateBench: Evaluating Moderation Models Agains…

200 篇论文

The rapid increase in hate speech on social media has exposed an unprecedented impact on society, making automated methods for detecting such content important. Unlike prior black-box models, we propose a novel transparent method for…

Hateful memes have become a significant concern on the Internet, necessitating robust automated detection systems. While Large Multimodal Models (LMMs) have shown promise in hateful meme detection, they face notable challenges like…

计算与语言 · 计算机科学 2026-03-03 Jingbiao Mei , Jinghong Chen , Guangyu Yang , Weizhe Lin , Bill Byrne

Detecting hateful content is a challenging and important problem. Automated tools, like machine-learning models, can help, but they require continuous training to adapt to the ever-changing landscape of social media. In this work, we…

计算与语言 · 计算机科学 2025-11-06 Jay Patel , Hrudayangam Mehta , Jeremy Blackburn

Generated hateful and toxic content by a portion of users in social media is a rising phenomenon that motivated researchers to dedicate substantial efforts to the challenging direction of hateful content identification. We not only need an…

社会与信息网络 · 计算机科学 2019-10-29 Marzieh Mozafari , Reza Farahbakhsh , Noel Crespi

In the past few years, there has been a significant rise in toxic and hateful content on various social media platforms. Recently Black Lives Matter movement came into the picture, causing an avalanche of user generated responses on the…

计算与语言 · 计算机科学 2021-08-31 Sumit Kumar , Raj Ratn Pranesh

Online hate remains a significant societal challenge, especially as multimodal content enables subtle, culturally grounded, and implicit forms of harm. Hateful memes embed hostility through text-image interactions and humor, making them…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Sahajpreet Singh , Kokil Jaidka , Subhayan Mukerjee

With the spread of social networks and their unfortunate use for hate speech, automatic detection of the latter has become a pressing problem. In this paper, we reproduce seven state-of-the-art hate speech detection models from prior work,…

计算与语言 · 计算机科学 2018-11-06 Tommi Gröndahl , Luca Pajola , Mika Juuti , Mauro Conti , N. Asokan

An increasingly common expression of online hate speech is multimodal in nature and comes in the form of memes. Designing systems to automatically detect hateful content is of paramount importance if we are to mitigate its undesirable…

Given the black-box nature and complexity of large transformer language models (LM), concerns about generalizability and robustness present ethical implications for domains such as hate speech (HS) detection. Using the content rich Social…

计算与语言 · 计算机科学 2024-11-12 Jennifer L. Chen , Faisal Ladhak , Daniel Li , Noémie Elhadad

Counterspeech can be an effective method for battling hateful content on social media. Automated counterspeech generation can aid in this process. Generated counterspeech, however, can be viable only when grounded in the context of topic,…

计算与语言 · 计算机科学 2023-12-01 Sabit Hassan , Malihe Alikhani

Dark humor often relies on subtle cultural nuances and implicit cues that require contextual reasoning to interpret, posing safety challenges that current static benchmarks fail to capture. To address this, we introduce a novel multimodal,…

计算与语言 · 计算机科学 2026-03-20 Ahmed Sharshar , Hosam Elgendy , Saad El Dine Ahmed , Yasser Rohaim , Yuxia Wang

Toxic language detection systems often falsely flag text that contains minority group mentions as toxic, as those groups are often the targets of online hate. Such over-reliance on spurious correlations also causes systems to struggle with…

计算与语言 · 计算机科学 2022-07-15 Thomas Hartvigsen , Saadia Gabriel , Hamid Palangi , Maarten Sap , Dipankar Ray , Ece Kamar

White supremacists embrace a radical ideology that considers white people superior to people of other races. The critical influence of these groups is no longer limited to social media; they also have a significant effect on society in many…

计算与语言 · 计算机科学 2020-10-02 Hind Saleh Alatawi , Areej Maatog Alhothali , Kawthar Mustafa Moria

Hate speech has emerged as a major problem plaguing our social spaces today. While there have been significant efforts to address this problem, existing methods are still significantly limited in effectively detecting hate speech online. A…

计算机与社会 · 计算机科学 2024-01-09 Keyan Guo , Alexander Hu , Jaden Mu , Ziheng Shi , Ziming Zhao , Nishant Vishwamitra , Hongxin Hu

Exploiting social media to spread hate has tremendously increased over the years. Lately, multi-modal hateful content such as memes has drawn relatively more traction than uni-modal content. Moreover, the availability of implicit content…

计算与语言 · 计算机科学 2023-02-14 Piush Aggarwal , Pranit Chawla , Mithun Das , Punyajoy Saha , Binny Mathew , Torsten Zesch , Animesh Mukherjee

Hateful and offensive content detection has been extensively explored in a single modality such as text. However, such toxic information could also be communicated via multimodal content such as online memes. Therefore, detecting multimodal…

信息检索 · 计算机科学 2021-08-16 Rui Cao , Ziqing Fan , Roy Ka-Wei Lee , Wen-Haw Chong , Jing Jiang

With the increasing diversity of use cases of large language models, a more informative treatment of texts seems necessary. An argumentative analysis could foster a more reasoned usage of chatbots, text completion mechanisms or other…

计算与语言 · 计算机科学 2023-06-06 Damián Furman , Pablo Torres , José A. Rodríguez , Diego Letzen , Vanina Martínez , Laura Alonso Alemany

Warning: This paper contains examples of the language that some people may find offensive. Detecting and reducing hateful, abusive, offensive comments is a critical and challenging task on social media. Moreover, few studies aim to mitigate…

计算与语言 · 计算机科学 2023-12-21 Neeraj Kumar Singh , Koyel Ghosh , Joy Mahapatra , Utpal Garain , Apurbalal Senapati

Hate speech represents a pervasive and detrimental form of online discourse, often manifested through an array of slurs, from hateful tweets to defamatory posts. As such speech proliferates, it connects people globally and poses significant…

计算与语言 · 计算机科学 2025-05-06 Paloma Piot , Patricia Martín-Rodilla , Javier Parapar

Recent advances in text-to-image diffusion models have enabled the creation of a new form of digital art: optical illusions--visual tricks that create different perceptions of reality. However, adversaries may misuse such techniques to…

密码学与安全 · 计算机科学 2025-07-31 Yiting Qu , Ziqing Yang , Yihan Ma , Michael Backes , Savvas Zannettou , Yang Zhang