中文
相关论文

相关论文: DeMod: A Holistic Tool with Explainable Detection …

200 篇论文

Content moderation faces a challenging task as social media's ability to spread hate speech contrasts with its role in promoting global connectivity. With rapidly evolving slang and hate speech, the adaptability of conventional deep…

机器学习 · 计算机科学 2024-04-18 Paras Sheth , Tharindu Kumarage , Raha Moraffah , Aman Chadha , Huan Liu

Regulatory bodies worldwide are intensifying their efforts to ensure transparency in influencer marketing on social media through instruments like the Unfair Commercial Practices Directive (UCPD) in the European Union, or Section 5 of the…

计算与语言 · 计算机科学 2023-06-09 Thales Bertaglia , Stefan Huber , Catalina Goanta , Gerasimos Spanakis , Adriana Iamnitchi

Background: The existence of toxic conversations in open-source platforms can degrade relationships among software developers and may negatively impact software product quality. To help mitigate this, some initial work has been done to…

软件工程 · 计算机科学 2023-07-10 Jaydeb Saker , Sayma Sultana , Steven R. Wilson , Amiangshu Bosu

Bridging content that brings together individuals with opposing viewpoints on social media remains elusive, overshadowed by echo chambers and toxic exchanges. We propose that algorithmic curation could surface such content by considering…

社会与信息网络 · 计算机科学 2025-09-24 Ozgur Can Seckin , Bao Tran Truong , Alessandro Flammini , Filippo Menczer

With surge in online platforms, there has been an upsurge in the user engagement on these platforms via comments and reactions. A large portion of such textual comments are abusive, rude and offensive to the audience. With machine learning…

计算与语言 · 计算机科学 2021-08-17 Ayush Kumar , Pratik Kumar

While recent research has focused on developing safeguards for generative AI (GAI) model-level content safety, little is known about how content moderation to prevent malicious content performs for end-users in real-world GAI products. To…

人机交互 · 计算机科学 2025-06-18 Lan Gao , Oscar Chen , Rachel Lee , Nick Feamster , Chenhao Tan , Marshini Chetty

The proliferation of online toxic speech is a pertinent problem posing threats to demographic groups. While explicit toxic speech contains offensive lexical signals, implicit one consists of coded or indirect language. Therefore, it is…

计算与语言 · 计算机科学 2024-05-21 Nhat M. Hoang , Xuan Long Do , Duc Anh Do , Duc Anh Vu , Luu Anh Tuan

Understanding toxicity in user conversations is undoubtedly an important problem. Addressing "covert" or implicit cases of toxicity is particularly hard and requires context. Very few previous studies have analysed the influence of…

计算与语言 · 计算机科学 2022-10-19 Atijit Anuchitanukul , Julia Ive , Lucia Specia

The mental health of social media users has started more and more to be put at risk by harmful, hateful, and offensive content. In this paper, we propose \textsc{StopHC}, a harmful content detection and mitigation architecture for social…

社会与信息网络 · 计算机科学 2024-11-12 Ciprian-Octavian Truică , Ana-Teodora Constantinescu , Elena-Simona Apostol

We present TruthBot, an all-in-one multilingual conversational chatbot designed for seeking truth (trustworthy and verified information) on specific topics. It helps users to obtain information specific to certain topics, fact-check…

社会与信息网络 · 计算机科学 2021-02-02 Ankur Gupta , Yash Varun , Prarthana Das , Nithya Muttineni , Parth Srivastava , Hamim Zafar , Tanmoy Chakraborty , Swaprava Nath

Toxic and antisocial user behavior on social media platforms has received considerable scholarly attention due to its detrimental effects on society. This study takes a holistic perspective on the phenomenon of online toxicity by…

社会与信息网络 · 计算机科学 2025-11-24 Lorenzo Alvisi , Victoria Popa , Guglielmo Cola , Serena Tardelli , Maurizio Tesconi

Software projects thrive on the involvement and contributions of individuals from different backgrounds. However, toxic language and negative interactions can hinder the participation and retention of contributors and alienate newcomers.…

In this paper, we assess existing technical proposals for content moderation in End-to-End Encryption (E2EE) services. First, we explain the various tools in the content moderation toolbox, how they are used, and the different phases of the…

密码学与安全 · 计算机科学 2022-02-10 Seny Kamara , Mallory Knodel , Emma Llansó , Greg Nojeim , Lucy Qin , Dhanaraj Thakur , Caitlin Vogus

The censorship of toxic comments is often left to the judgment of imperfect models. Perspective API, a creation of Google technology incubator Jigsaw, is perhaps the most widely used toxicity classifier in industry; the model is employed by…

计算机与社会 · 计算机科学 2020-07-01 Elizabeth Reichert , Helen Qiu , Jasmine Bayrooti

Content moderation (removing or limiting the distribution of posts based on their contents) is one tool social networks use to fight problems such as harassment and disinformation. Manually screening all content is usually impractical given…

信息检索 · 计算机科学 2021-08-31 Eugene Yang , David D. Lewis , Ophir Frieder

The detection of suicide risk in social media is a critical task with potential life-saving implications. This paper presents a study on leveraging state-of-the-art natural language processing solutions for identifying suicide risk in…

计算与语言 · 计算机科学 2024-10-14 Jakub Pokrywka , Jeremi I. Kaczmarek , Edward J. Gorzelańczyk

We present Moderator, a policy-based model management system that allows administrators to specify fine-grained content moderation policies and modify the weights of a text-to-image (TTI) model to make it significantly more challenging for…

密码学与安全 · 计算机科学 2024-09-13 Peiran Wang , Qiyu Li , Longxuan Yu , Ziyao Wang , Ang Li , Haojian Jin

Debiasing methods that seek to mitigate the tendency of Language Models (LMs) to occasionally output toxic or inappropriate text have recently gained traction. In this paper, we propose a standardized protocol which distinguishes methods…

计算与语言 · 计算机科学 2023-05-24 Robert Morabito , Jad Kabbara , Ali Emami

The proliferation of generative models has presented significant challenges in distinguishing authentic human-authored content from deepfake content. Collaborative human efforts, augmented by AI tools, present a promising solution. In this…

计算与语言 · 计算机科学 2026-03-25 Jooyoung Lee , Xiaochen Zhu , Georgi Karadzhov , Tom Stafford , Andreas Vlachos , Dongwon Lee

Biased associations have been a challenge in the development of classifiers for detecting toxic language, hindering both fairness and accuracy. As potential solutions, we investigate recently introduced debiasing methods for text…

计算与语言 · 计算机科学 2021-02-02 Xuhui Zhou , Maarten Sap , Swabha Swayamdipta , Noah A. Smith , Yejin Choi