中文
相关论文

相关论文: Explainable Abuse Detection as Intent Classificati…

200 篇论文

Content safety teams need metrics that reflect what users actually experience, not only what is reported. We study prevalence: the fraction of user views (impressions) that went to content violating a given policy on a given day. Accurate…

机器学习 · 计算机科学 2026-02-24 Attila Dobi , Aravindh Manickavasagam , Benjamin Thompson , Xiaohan Yang , Faisal Farooq

Fairness of machine learning models in healthcare has drawn increasing attention from clinicians, researchers, and even at the highest level of government. On the other hand, the importance of developing and deploying interpretable or…

The proliferation of social media platforms and online communities has inadvertently catalyzed the spread of cyberbullying, hate speech, and other forms of online toxicity, making the effective governance of such harm a critical societal…

人工智能 · 计算机科学 2026-05-28 Yiting Huang , Wenting Zhu , Zekun Wang , Qingpo Yang , Yakai Chen , Zihui Xu , Yueyue Zhang , Sanchuan Guo , Xi Zhang

Abuse on the Internet represents an important societal problem of our time. Millions of Internet users face harassment, racism, personal attacks, and other types of abuse on online platforms. The psychological effects of such abuse on…

计算与语言 · 计算机科学 2020-10-01 Pushkar Mishra , Helen Yannakoudakis , Ekaterina Shutova

Fringe groups and organizations have a long history of using euphemisms--ordinary-sounding words with a secret meaning--to conceal what they are discussing. Nowadays, one common use of euphemisms is to evade content moderation policies…

计算与语言 · 计算机科学 2021-04-01 Wanzheng Zhu , Hongyu Gong , Rohan Bansal , Zachary Weinberg , Nicolas Christin , Giulia Fanti , Suma Bhat

Social networks have become an increasingly common abstraction to capture the interactions of individual users in a number of everyday activities and applications. As a result, the analysis of such networks has attracted lots of attention…

社会与信息网络 · 计算机科学 2023-05-05 Ahmad Zareie , Rizos Sakellariou

An-ever increasing number of social media websites, electronic newspapers and Internet forums allow visitors to leave comments for others to read and interact. This exchange is not free from participants with malicious intentions, which do…

计算与语言 · 计算机科学 2017-04-11 Luis Gerardo Mojica

This paper presents a pipeline to detect and explain anomalous reviews in online platforms. The pipeline is made up of three modules and allows the detection of reviews that do not generate value for users due to either worthless or…

计算与语言 · 计算机科学 2024-02-29 David Novoa-Paradela , Oscar Fontenla-Romero , Bertha Guijarro-Berdiñas

Extensive efforts in automated approaches for content moderation have been focused on developing models to identify toxic, offensive, and hateful content with the aim of lightening the load for moderators. Yet, it remains uncertain whether…

计算与语言 · 计算机科学 2024-11-14 Yang Trista Cao , Lovely-Frances Domingo , Sarah Ann Gilbert , Michelle Mazurek , Katie Shilton , Hal Daumé

Urban transit agencies increasingly turn to social media to monitor emerging service risks such as crowding, delays, and safety incidents, yet the signals of concern are sparse, short, and easily drowned by routine chatter. We address this…

机器学习 · 计算机科学 2025-12-09 Fatima Ashraf , Muhammad Ayub Sabir , Jiaxin Deng , Junbiao Pang , Haitao Yu

Technologies for abusive language detection are being developed and applied with little consideration of their potential biases. We examine racial bias in five different sets of Twitter data annotated for hate speech and abusive language.…

计算与语言 · 计算机科学 2019-05-30 Thomas Davidson , Debasmita Bhattacharya , Ingmar Weber

The proliferation of harmful content on online platforms is a major societal problem, which comes in many different forms including hate speech, offensive language, bullying and harassment, misinformation, spam, violence, graphic content,…

Current content moderation follows a reactive, trial-and-error approach, where interventions are applied and their effects are only measured post-hoc. In contrast, we introduce a proactive, predictive approach that enables moderators to…

计算机与社会 · 计算机科学 2026-02-09 Benedetta Tessa , Lorenzo Cima , Amaury Trujillo , Marco Avvenuti , Stefano Cresci

The detection of sensitive content in large datasets is crucial for ensuring that shared and analysed data is free from harmful material. However, current moderation tools, such as external APIs, suffer from limitations in customisation,…

计算与语言 · 计算机科学 2025-06-25 Dimosthenis Antypas , Indira Sen , Carla Perez-Almendros , Jose Camacho-Collados , Francesco Barbieri

Many real incidents demonstrate that users of Online Social Networks need mechanisms that help them manage their interactions by increasing the awareness of the different contexts that coexist in Online Social Networks and preventing them…

社会与信息网络 · 计算机科学 2016-06-14 Natalia Criado , Jose M. Such

Evaluating the safety of AI Systems is a pressing concern for organizations deploying them. In addition to the societal damage done by the lack of fairness of those systems, deployers are concerned about the legal repercussions and the…

Content moderation is a widely used strategy to prevent the dissemination of irregular information on social media platforms. Despite extensive research on developing automated models to support decision-making in content moderation, there…

社会与信息网络 · 计算机科学 2024-08-23 Wangjiaxuan Xin , Kanlun Wang , Zhe Fu , Lina Zhou

The importance of social media in our daily lives has unfortunately led to an increase in the spread of misinformation, political messages and malicious links. One of the most popular ways of carrying out those activities is using automated…

社会与信息网络 · 计算机科学 2024-11-12 Salvador Lopez-Joya , Jose A. Diaz-Garcia , M. Dolores Ruiz , Maria J. Martin-Bautista

The rise of social media has been argued to intensify uncivil and hostile online political discourse. Yet, to date, there is a lack of clarity on what incivility means in the political sphere. In this work, we utilize a multidimensional…

计算与语言 · 计算机科学 2023-11-16 Sagi Pendzel , Nir Lotan , Alon Zoizner , Einat Minkov

The abstract outlines the problem of toxic comments on social media platforms, where individuals use disrespectful, abusive, and unreasonable language that can drive users away from discussions. This behavior is referred to as anti-social…

机器学习 · 计算机科学 2023-04-17 K. Poojitha , A. Sai Charish , M. Arun Kuamr Reddy , S. Ayyasamy