中文
相关论文

相关论文: Improving Implicit Hate Speech Detection via a Com…

200 篇论文

In recent years, hate speech has gained great relevance in social networks and other virtual media because of its intensity and its relationship with violent acts against members of protected groups. Due to the great amount of content…

Considering the importance of detecting hateful language, labeled hate speech data is expensive and time-consuming to collect, particularly for low-resource languages. Prior work has demonstrated the effectiveness of cross-lingual transfer…

计算与语言 · 计算机科学 2025-05-27 Faeze Ghorbanpour , Daryna Dementieva , Alexander Fraser

While significant progress has been made using machine learning algorithms to detect hate speech, important technical challenges still remain to be solved in order to bring their performance closer to human accuracy. We investigate several…

机器学习 · 计算机科学 2020-12-25 Vlad Sandulescu

Employing language models to generate explanations for an incoming implicit hate post is an active area of research. The explanation is intended to make explicit the underlying stereotype and aid content moderators. The training often…

计算与语言 · 计算机科学 2024-06-07 Neemesh Yadav , Sarah Masud , Vikram Goyal , Vikram Goyal , Md Shad Akhtar , Tanmoy Chakraborty

The context-dependent nature of online aggression makes annotating large collections of data extremely difficult. Previously studied datasets in abusive language detection have been insufficient in size to efficiently train deep learning…

计算与语言 · 计算机科学 2018-08-31 Younghun Lee , Seunghyun Yoon , Kyomin Jung

Today, the internet is an integral part of our daily lives, enabling people to be more connected than ever before. However, this greater connectivity and access to information increase exposure to harmful content such as cyber-bullying and…

社会与信息网络 · 计算机科学 2023-10-31 Lanqin Yuan , Tianyu Wang , Gabriela Ferraro , Hanna Suominen , Marian-Andrei Rizoiu

Hate speech detection is commonly framed as a direct binary classification problem despite being a composite concept defined through multiple interacting factors that vary across legal frameworks, platform policies, and annotation…

计算与语言 · 计算机科学 2026-02-06 Adrián Girón , Pablo Miralles , Javier Huertas-Tato , Sergio D'Antonio , David Camacho

Automatic detection of toxic language plays an essential role in protecting social media users, especially minority groups, from verbal abuse. However, biases toward some attributes, including gender, race, and dialect, exist in most…

计算与语言 · 计算机科学 2021-06-15 Yung-Sung Chuang , Mingye Gao , Hongyin Luo , James Glass , Hung-yi Lee , Yun-Nung Chen , Shang-Wen Li

Hate, derogatory, and offensive speech remains a persistent challenge in online platforms and public discourse. While automated detection systems are widely used, most focus on censorship or removal, raising concerns for transparency and…

With the proliferation of social media, there has been a sharp increase in offensive content, particularly targeting vulnerable groups, exacerbating social problems such as hatred, racism, and sexism. Detecting offensive language use is…

计算与语言 · 计算机科学 2023-12-05 Toygar Tanyel , Besher Alkurdi , Serkan Ayvaz

We present a novel feature attribution method for explaining text classifiers, and analyze it in the context of hate speech detection. Although feature attribution models usually provide a single importance score for each token, we instead…

计算与语言 · 计算机科学 2022-05-09 Esma Balkir , Isar Nejadgholi , Kathleen C. Fraser , Svetlana Kiritchenko

Automatic detection of online hate speech serves as a crucial step in the detoxification of the online discourse. Moreover, accurate classification can promote a better understanding of the proliferation of hate as a social phenomenon.…

计算与语言 · 计算机科学 2025-06-25 Tom Marzea , Abraham Israeli , Oren Tsur

With the proliferation of social media, accurate detection of hate speech has become critical to ensure safety online. To combat nuanced forms of hate speech, it is important to identify and thoroughly explain hate speech to help users…

计算与语言 · 计算机科学 2023-11-23 Yongjin Yang , Joonkee Kim , Yujin Kim , Namgyu Ho , James Thorne , Se-young Yun

Threat detection in Natural Language Processing lacks consistent definitions and standardized benchmarks, and is often conflated with broader phenomena such as toxicity, hate speech, or offensive language. In this work, we introduce…

计算与语言 · 计算机科学 2026-05-12 Davide Bruni , Carlo Bardazzi , Maurizio Tesconi

Large language model (LLM) agents increasingly rely on external tools and retrieval systems to autonomously complete complex tasks. However, this design exposes agents to indirect prompt injection (IPI), where attacker-controlled context…

密码学与安全 · 计算机科学 2026-02-27 Tian Zhang , Yiwei Xu , Juan Wang , Keyan Guo , Xiaoyang Xu , Bowen Xiao , Quanlong Guan , Jinlin Fan , Jiawei Liu , Zhiquan Liu , Hongxin Hu

We present Conformal Intent Classification and Clarification (CICC), a framework for fast and accurate intent classification for task-oriented dialogue systems. The framework turns heuristic uncertainty scores of any intent classifier into…

计算与语言 · 计算机科学 2024-03-29 Floris den Hengst , Ralf Wolter , Patrick Altmeyer , Arda Kaygan

The proliferation of hate speech has inflicted significant societal harm, with its intensity and directionality closely tied to specific targets and arguments. In recent years, numerous machine learning-based methods have been developed to…

计算与语言 · 计算机科学 2025-07-16 Zewen Bai , Liang Yang , Shengdi Yin , Yuanyuan Sun , Hongfei Lin

Hate speech, offensive language, aggression, racism, sexism, and other abusive language are common phenomena in social media. There is a need for Artificial Intelligence(AI)based intervention which can filter hate content at scale. Most…

计算与语言 · 计算机科学 2024-11-13 Prashant Kapil , Asif Ekbal

We introduce AgenticSimLaw, a role-structured, multi-agent debate framework that provides transparent and controllable test-time reasoning for high-stakes tabular decision-making tasks. Unlike black-box approaches, our courtroom-style…

人工智能 · 计算机科学 2026-01-30 Jon Chun , Kathrine Elkins , Yong Suk Lee

Hate speech is one type of harmful online content which directly attacks or promotes hate towards a group or an individual member based on their actual or perceived aspects of identity, such as ethnicity, religion, and sexual orientation.…

计算与语言 · 计算机科学 2021-02-18 Wenjie Yin , Arkaitz Zubiaga