中文
相关论文

相关论文: Designing Toxic Content Classification for a Diver…

200 篇论文

Social network platforms are generally used to share positive, constructive, and insightful content. However, in recent times, people often get exposed to objectionable content like threat, identity attacks, hate speech, insults, obscene…

计算与语言 · 计算机科学 2021-05-31 Sreyan Ghosh , Sonal Kumar

Perception of toxicity evolves over time and often differs between geographies and cultural backgrounds. Similarly, black-box commercially available APIs for detecting toxicity, such as the Perspective API, are not static, but frequently…

计算与语言 · 计算机科学 2023-04-26 Luiza Pozzobon , Beyza Ermis , Patrick Lewis , Sara Hooker

Online debates are often characterised by extreme polarisation and heated discussions among users. The presence of hate speech online is becoming increasingly problematic, making necessary the development of appropriate countermeasures. In…

社会与信息网络 · 计算机科学 2021-05-31 Matteo Cinelli , Andraž Pelicon , Igor Mozetič , Walter Quattrociocchi , Petra Kralj Novak , Fabiana Zollo

Social media platforms provide users the freedom of expression and a medium to exchange information and express diverse opinions. Unfortunately, this has also resulted in the growth of abusive content with the purpose of discriminating…

计算与语言 · 计算机科学 2021-07-01 Sohail Akhtar , Valerio Basile , Viviana Patti

People with disabilities (PwD) regularly encounter ableist hate and microaggressions online. These spaces are generally moderated by machine learning models, but little is known about how effectively AI models identify ableist speech and…

人机交互 · 计算机科学 2025-06-03 Mahika Phutane , Ananya Seelam , Aditya Vashistha

AI systems crucially rely on human ratings, but these ratings are often aggregated, obscuring the inherent diversity of perspectives in real-world phenomenon. This is particularly concerning when evaluating the safety of generative AI,…

The rapid progress of generative AI has enabled remarkable creative capabilities, yet it also raises urgent concerns regarding the safety of AI-generated visual content in real-world applications such as content moderation, platform…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Qiang Fu , Zonglei Jing , Zonghao Ying , Xiaoqian Li

The convenience of social media has also enabled its misuse, potentially resulting in toxic behavior. Nearly 66% of internet users have observed online harassment, and 41% claim personal experience, with 18% facing severe forms of online…

Toxic language, such as hate speech, can deter users from participating in online communities and enjoying popular platforms. Previous approaches to detecting toxic language and norm violations have been primarily concerned with…

计算与语言 · 计算机科学 2023-10-10 Jihyung Moon , Dong-Ho Lee , Hyundong Cho , Woojeong Jin , Chan Young Park , Minwoo Kim , Jonathan May , Jay Pujara , Sungjoon Park

The presence of toxic and gender-identity derogatory language in open-source software (OSS) communities has recently become a focal point for researchers. Such comments not only lead to frustration and disengagement among developers but may…

软件工程 · 计算机科学 2024-03-18 Sayma Sultana , Gias Uddin , Amiangshu Bosu

Although automated harmful content detection systems are frequently used to monitor online platforms, moderators and end users frequently cannot understand the logic underlying their predictions. While recent studies have focused on…

计算与语言 · 计算机科学 2026-03-20 Trishita Dhara , Siddhesh Sheth

The automatic identification of harmful content online is of major concern for social media platforms, policymakers, and society. Researchers have studied textual, visual, and audio content, but typically in isolation. Yet, harmful content…

The ability to accurately detect and filter offensive content automatically is important to ensure a rich and diverse digital discourse. Trolling is a type of hurtful or offensive content that is prevalent in social media, but is…

计算机与社会 · 计算机科学 2020-08-04 Hitkul , Karmanya Aggarwal , Pakhi Bamdev , Debanjan Mahata , Rajiv Ratn Shah , Ponnurangam Kumaraguru

In this work, we examine the influence of unreliable information on political incivility and toxicity on the social media platform Reddit. We show that comments on articles from unreliable news websites are posted more often in…

社会与信息网络 · 计算机科学 2024-11-01 Hans W. A. Hanley , Zakir Durumeric

Harmful content detection models tend to have higher false positive rates for content from marginalized groups. In the context of marginal abuse modeling on Twitter, such disproportionate penalization poses the risk of reduced visibility,…

计算与语言 · 计算机科学 2022-10-13 Kyra Yee , Alice Schoenauer Sebag , Olivia Redfield , Emily Sheng , Matthias Eck , Luca Belli

Now that AI-driven moderation has become pervasive in everyday life, we often hear claims that "the AI is biased". While this is often said jokingly, the light-hearted remark reflects a deeper concern. How can we be certain that an online…

计算与语言 · 计算机科学 2026-04-02 Subhojit Ghimire

We analyze sentiment analysis and toxicity detection models to detect the presence of explicit bias against people with disability (PWD). We employ the bias identification framework of Perturbation Sensitivity Analysis to examine…

计算与语言 · 计算机科学 2023-07-19 Pranav Narayanan Venkit , Mukund Srinath , Shomir Wilson

Toxicity is endemic to online social networks including Twitter. It follows a Pareto like distribution where most of the toxicity is generated by a very small number of profiles and as such, analyzing and characterizing these toxic profiles…

社会与信息网络 · 计算机科学 2023-03-28 Hina Qayyum , Benjamin Zi Hao Zhao , Ian D. Wood , Muhammad Ikram , Mohamed Ali Kaafar , Nicolas Kourtellis

While the contagious nature of online toxicity sparked increasing interest in its early detection and prevention, most of the literature focuses on the Western world. In this work, we demonstrate that 1) it is possible to detect toxicity…

计算机与社会 · 计算机科学 2022-04-20 Yun Yu Chong , Haewoon Kwak

Toxicity has become a grave problem for many online communities and has been growing across many languages, including Russian. Hate speech creates an environment of intimidation, discrimination, and may even incite some real-world violence.…

计算与语言 · 计算机科学 2020-10-23 Nadezhda Zueva , Madina Kabirova , Pavel Kalaidin