English
Related papers

Related papers: Detecting Inappropriate Messages on Sensitive Topi…

200 papers

Online social platforms are beset with hateful speech - content that expresses hatred for a person or group of people. Such content can frighten, intimidate, or silence platform users, and some of it can inspire other users to commit…

Computation and Language · Computer Science 2017-10-02 Haji Mohammad Saleem , Kelly P Dillon , Susan Benesch , Derek Ruths

We present a holistic approach to building a robust and useful natural language classification system for real-world content moderation. The success of such a system relies on a chain of carefully designed and executed steps, including the…

Computation and Language · Computer Science 2023-02-16 Todor Markov , Chong Zhang , Sandhini Agarwal , Tyna Eloundou , Teddy Lee , Steven Adler , Angela Jiang , Lilian Weng

Text toxicity detection systems exhibit significant biases, producing disproportionate rates of false positives on samples mentioning demographic groups. But what about toxicity detection in speech? To investigate the extent to which…

Online reviews provide viewpoints on the strengths and shortcomings of products/services, influencing potential customers' purchasing decisions. However, the proliferation of non-credible reviews -- either fake (promoting/ demoting an…

Artificial Intelligence · Computer Science 2017-05-09 Subhabrata Mukherjee , Sourav Dutta , Gerhard Weikum

To foster collaboration and inclusivity in Open Source Software (OSS) projects, it is crucial to understand and detect patterns of toxic language that may drive contributors away, especially those from underrepresented communities. Although…

Software Engineering · Computer Science 2023-07-31 Ramtin Ehsani , Rezvaneh Rezapour , Preetha Chatterjee

Models trained on large unlabeled corpora of human interactions will learn patterns and mimic behaviors therein, which include offensive or otherwise toxic behavior and unwanted biases. We investigate a variety of methods to mitigate these…

Computation and Language · Computer Science 2021-08-06 Jing Xu , Da Ju , Margaret Li , Y-Lan Boureau , Jason Weston , Emily Dinan

Perception of offensiveness is inherently subjective, shaped by the lived experiences and socio-cultural values of the perceivers. Recent years have seen substantial efforts to build AI-based tools that can detect offensive language at…

Computers and Society · Computer Science 2023-12-13 Aida Davani , Mark Díaz , Dylan Baker , Vinodkumar Prabhakaran

Detecting prosociality in text--communication intended to affirm, support, or improve others' behavior--is a novel and increasingly important challenge for trust and safety systems. Unlike toxic content detection, prosociality lacks…

Computation and Language · Computer Science 2025-08-11 Rafal Kocielnik , Min Kim , Penphob , Boonyarungsrit , Fereshteh Soltani , Deshawn Sambrano , Animashree Anandkumar , R. Michael Alvarez

Recent years have seen increasing concerns about the unsafe response generation of large-scale dialogue systems, where agents will learn offensive or biased behaviors from the real-world corpus. Some methods are proposed to address the…

Computation and Language · Computer Science 2023-05-26 Zi Liang , Pinghui Wang , Ruofei Zhang , Shuo Zhang , Xiaofan Ye Yi Huang , Junlan Feng

While deep learning models have greatly improved the performance of most artificial intelligence tasks, they are often criticized to be untrustworthy due to the black-box problem. Consequently, many works have been proposed to study the…

Computation and Language · Computer Science 2021-09-08 Lijie Wang , Hao Liu , Shuyuan Peng , Hongxuan Tang , Xinyan Xiao , Ying Chen , Hua Wu , Haifeng Wang

Topic models are valuable for understanding extensive document collections, but they don't always identify the most relevant topics. Classical probabilistic and anchor-based topic models offer interactive versions that allow users to guide…

Machine Learning · Computer Science 2024-02-08 Kyle Seelman , Mozhi Zhang , Jordan Boyd-Graber

The automatic identification of harmful content online is of major concern for social media platforms, policymakers, and society. Researchers have studied textual, visual, and audio content, but typically in isolation. Yet, harmful content…

Reputation is a central element of social communications, be it with human or artificial intelligence (AI), and as such can be the primary target of malicious communication strategies. There is already a vast amount of literature on trust…

Physics and Society · Physics 2022-05-18 Torsten Enßlin , Viktoria Kainz , Céline Bœhm

Researchers and developers increasingly rely on toxicity scoring to moderate generative language model outputs, in settings such as customer service, information retrieval, and content generation. However, toxicity scoring may render…

Human-Computer Interaction · Computer Science 2024-04-23 Jennifer Chien , Kevin R. McKee , Jackie Kay , William Isaac

We use structural topic modeling to examine racial bias in data collected to train models to detect hate speech and abusive language in social media posts. We augment the abusive language dataset by adding an additional feature indicating…

Computation and Language · Computer Science 2020-05-28 Thomas Davidson , Debasmita Bhattacharya

Telegram has become a major space for political discourse and alternative media. However, its lack of moderation allows misinformation, extremism, and toxicity to spread. While prior research focused on these particular phenomena or topics,…

Social and Information Networks · Computer Science 2025-05-06 Lorenzo Alvisi , Serena Tardelli , Maurizio Tesconi

With the recent proliferation of the use of text classifications, researchers have found that there are certain unintended biases in text classification datasets. For example, texts containing some demographic identity-terms (e.g., "gay",…

Computation and Language · Computer Science 2020-08-21 Guanhua Zhang , Bing Bai , Junqi Zhang , Kun Bai , Conghui Zhu , Tiejun Zhao

Studies have shown that toxic behavior can cause contributors to leave, and hinder newcomers' (especially from underrepresented communities) participation in Open Source Software (OSS) projects. Thus, detection of toxic language plays a…

Software Engineering · Computer Science 2025-01-28 Ramtin Ehsani , Rezvaneh Rezapour , Preetha Chatterjee

Now that AI-driven moderation has become pervasive in everyday life, we often hear claims that "the AI is biased". While this is often said jokingly, the light-hearted remark reflects a deeper concern. How can we be certain that an online…

Computation and Language · Computer Science 2026-04-02 Subhojit Ghimire

Online discussions, panels, talk page edits, etc., often contain harmful conversational content i.e., hate speech, death threats and offensive language, especially towards certain demographic groups. For example, individuals who identify as…

Computation and Language · Computer Science 2022-07-21 Jamell Dacon , Harry Shomer , Shaylynn Crum-Dacon , Jiliang Tang