中文
相关论文

相关论文: DeMod: A Holistic Tool with Explainable Detection …

200 篇论文

Personal moderation tools on social media platforms let users control their news feeds by configuring acceptable toxicity thresholds for their feed content or muting inappropriate accounts. This research examines how four critical…

人机交互 · 计算机科学 2024-03-05 Shagun Jhaver

Platforms that support online commentary, from social networks to news sites, are increasingly leveraging machine learning to assist their moderation efforts. But this process does not typically provide feedback to the author that would…

计算与语言 · 计算机科学 2021-02-12 Leo Laugier , John Pavlopoulos , Jeffrey Sorensen , Lucas Dixon

As toxic language becomes nearly pervasive online, there has been increasing interest in leveraging the advancements in natural language processing (NLP), from very large transformer models to automatically detecting and removing toxic…

计算与语言 · 计算机科学 2020-07-02 Austin P. Wright , Omar Shaikh , Haekyu Park , Will Epperson , Muhammed Ahmed , Stephane Pinel , Diyi Yang , Duen Horng Chau

In the digital world, memes present a unique challenge for content moderation due to their potential to spread harmful content. Although detection methods have improved, proactive solutions such as intervention are still limited, with…

计算与语言 · 计算机科学 2024-06-11 Prince Jha , Raghav Jain , Konika Mandal , Aman Chadha , Sriparna Saha , Pushpak Bhattacharyya

Language models (LMs) can reproduce (or amplify) toxic language seen during training, which poses a risk to their practical application. In this paper, we conduct extensive experiments to study this phenomenon. We analyze the impact of…

计算与语言 · 计算机科学 2022-03-08 Canwen Xu , Zexue He , Zhankui He , Julian McAuley

In this work, we demonstrate how existing classifiers for identifying toxic comments online fail to generalize to the diverse concerns of Internet users. We survey 17,280 participants to understand how user expectations for what constitutes…

社会与信息网络 · 计算机科学 2021-06-09 Deepak Kumar , Patrick Gage Kelley , Sunny Consolvo , Joshua Mason , Elie Bursztein , Zakir Durumeric , Kurt Thomas , Michael Bailey

Social media platforms struggle to protect users from harmful content through content moderation. These platforms have recently leveraged machine learning models to cope with the vast amount of user-generated content daily. Since moderation…

机器学习 · 计算机科学 2023-01-27 Donghyun Son , Byounggyu Lew , Kwanghee Choi , Yongsu Baek , Seungwoo Choi , Beomjun Shin , Sungjoo Ha , Buru Chang

Existing studies have investigated the tendency of autoregressive language models to generate contexts that exhibit undesired biases and toxicity. Various debiasing approaches have been proposed, which are primarily categorized into…

计算与语言 · 计算机科学 2022-05-03 Yoon A Park , Frank Rudzicz

Toxic content detection aims to identify content that can offend or harm its recipients. Automated classifiers of toxic content need to be robust against adversaries who deliberately try to bypass filters. We propose a method of generating…

计算与语言 · 计算机科学 2019-12-17 Keita Kurita , Anna Belova , Antonios Anastasopoulos

The emergence of social media has largely eased the way people receive information and participate in public discussions. However, in countries with strict regulations on discussions in the public space, social media is no exception. To…

社会与信息网络 · 计算机科学 2022-07-26 Yichi Qian , Qiyi Shan , Hanjia Lyu , Jiebo Luo

Online abuse is becoming an increasingly prevalent issue in modern-day society, with 41 percent of Americans having experienced online harassment in some capacity in 2021. People who identify as women, in particular, can be subjected to a…

人机交互 · 计算机科学 2023-01-19 Sarah Barrington

Recent advances in natural language processing and machine learning have led to the development of chatbot models, such as ChatGPT, that can engage in conversational dialogue with human users. However, the ability of these models to…

密码学与安全 · 计算机科学 2023-07-20 Bocheng Chen , Guangjing Wang , Hanqing Guo , Yuanda Wang , Qiben Yan

The rapid growth of social media platforms has raised significant concerns regarding online content toxicity. When Large Language Models (LLMs) are used for toxicity detection, two key challenges emerge: 1) the absence of domain-specific…

计算与语言 · 计算机科学 2025-06-03 Yibo Zhao , Jiapeng Zhu , Can Xu , Yao Liu , Xiang Li

Decentralized communication is becoming an important use case within Web3. On Ethereum, users can repurpose the transaction input data field to embed natural-language messages, commonly known as Input Data Messages (IDMs). However, as IDMs…

密码学与安全 · 计算机科学 2025-10-14 Xihan Xiong , Zhipeng Wang , Qin Wang , William Knottenbelt

Toxicity detection is crucial for maintaining the peace of the society. While existing methods perform well on normal toxic contents or those generated by specific perturbation methods, they are vulnerable to evolving perturbation patterns.…

密码学与安全 · 计算机科学 2025-03-05 Hankun Kang , Jianhao Chen , Yongqi Li , Xin Miao , Mayi Xu , Ming Zhong , Yuanyuan Zhu , Tieyun Qian

Large language models (LLMs) exhibit impressive language capabilities but remain vulnerable to malicious prompts and jailbreaking attacks. Existing knowledge editing methods for LLM detoxification face two major challenges. First, they…

计算与语言 · 计算机科学 2025-05-29 Yifan Lu , Jing Li , Yigeng Zhou , Yihui Zhang , Wenya Wang , Xiucheng Li , Meishan Zhang , Fangming Liu , Jun Yu , Min Zhang

In the pursuit of bolstering user safety, social media platforms deploy active moderation strategies, including content removal and user suspension. These measures target users engaged in discussions marked by hate speech or toxicity, often…

社会与信息网络 · 计算机科学 2024-01-26 Hina Qayyum , Muhammad Ikram , Benjamin Zi Hao Zhao , Ian D. Wood , Nicolas Kourtellis , Mohamed Ali Kaafar

The training of large language models (LLMs) on extensive, unfiltered corpora sourced from the internet is a common and advantageous practice. Consequently, LLMs have learned and inadvertently reproduced various types of biases, including…

计算与语言 · 计算机科学 2023-11-20 Ambri Ma , Arnav Kumar , Brett Zeligson

Toxicity detection mitigates the dissemination of toxic content (e.g., hateful comments, posts, and messages within online social actions) to safeguard a healthy online social environment. However, malicious users persistently develop…

计算与语言 · 计算机科学 2026-03-17 Hankun Kang , Xin Miao , Jianhao Chen , Jintao Wen , Mayi Xu , Weiyu Zhang , Wenpeng Lu , Tieyun Qian

User-generated content (UGC) on social media platforms is vulnerable to incitements and manipulations, necessitating effective regulations. To address these challenges, those platforms often deploy automated content moderators tasked with…

机器学习 · 计算机科学 2025-07-29 Saba Ahmadi , Avrim Blum , Haifeng Xu , Fan Yao