中文
相关论文

相关论文: Harmful Speech Detection by Language Models Exhibi…

200 篇论文

Sensitive attributes are legally protected characteristics that should not be used to discriminate. Careful steps have been taken to minimize the risk of human bias regarding these fields, such as race and age. Large language models (LLMs)…

计算机与社会 · 计算机科学 2026-04-14 Anay Agarwalla , Simeon Sayer

Content moderation plays a critical role in shaping safe and inclusive online environments, balancing platform standards, user expectations, and regulatory frameworks. Traditionally, this process involves operationalising policies into…

Hateful comments are prevalent on social media platforms. Although tools for automatically detecting, flagging, and blocking such false, offensive, and harmful content online have lately matured, such reactive and brute force methods alone…

计算与语言 · 计算机科学 2024-01-17 Sougata Saha , Rohini Srihari

Detecting and classifying instances of hate in social media text has been a problem of interest in Natural Language Processing in the recent years. Our work leverages state of the art Transformer language models to identify hate speech in a…

计算与语言 · 计算机科学 2021-01-12 Sayar Ghosh Roy , Ujwal Narayan , Tathagata Raha , Zubair Abid , Vasudeva Varma

Eating disorders (ED), a severe mental health condition with high rates of mortality and morbidity, affect millions of people globally, especially adolescents. The proliferation of online communities that promote and normalize ED has been…

社会与信息网络 · 计算机科学 2024-05-24 Minh Duc Chu , Zihao He , Rebecca Dorn , Kristina Lerman

The rise of social media has significantly increased the prevalence of cyberbullying (CB), posing serious risks to both mental and physical well-being. Effective detection systems are essential for mitigating its impact. While several…

机器学习 · 计算机科学 2025-04-16 Tabia Tanzin Prama , Jannatul Ferdaws Amrin , Md. Mushfique Anwar , Iqbal H. Sarker

Transgender and non-binary (TGNB) individuals disproportionately experience discrimination and exclusion from daily life. Given the recent popularity and adoption of language generation technologies, the potential to further marginalize…

While social media offers freedom of self-expression, abusive language carry significant negative social impact. Driven by the importance of the issue, research in the automated detection of abusive language has witnessed growth and…

计算与语言 · 计算机科学 2022-05-04 Wenjie Yin , Arkaitz Zubiaga

Content moderation research has recently made significant advances, but remains limited in serving the majority of the world's languages due to the lack of resources, leaving millions of vulnerable users to online hostility. This work…

计算与语言 · 计算机科学 2025-10-28 Fitsum Gaim , Hoyun Song , Huije Lee , Changgeon Ko , Eui Jun Hwang , Jong C. Park

In today's digital world, social media plays a significant role in facilitating communication and content sharing. However, the exponential rise in user-generated content has led to challenges in maintaining a respectful online environment.…

计算与语言 · 计算机科学 2024-03-05 Mohammad Dehghani

As the application of Large Language Models (LLMs) spreads across various industries, there are increasing concerns about the potential for their misuse, especially in sensitive areas such as political discourse. Deliberately aligning LLMs…

计算与语言 · 计算机科学 2026-04-28 Lisa Korver , Mohamed Mostagir , Sherief Reda

Language models trained on large-scale unfiltered datasets curated from the open web acquire systemic biases, prejudices, and harmful views from their training data. We present a methodology for programmatically identifying and removing…

计算与语言 · 计算机科学 2021-11-30 Helen Ngo , Cooper Raterink , João G. M. Araújo , Ivan Zhang , Carol Chen , Adrien Morisot , Nicholas Frosst

This paper presents a multi-stage framework for detecting reclaimed slurs in multilingual social media discourse. It addresses the challenge of identifying reclamatory versus non-reclamatory usage of LGBTQ+-related slurs across English,…

计算与语言 · 计算机科学 2026-05-19 Barathi Ganesh HB , Michal Ptaszynski , Rene Melendez , Juuso Eronen

Online Social Networks (OSNs) widely adopt content moderation to mitigate the spread of abusive and toxic discourse. Nonetheless, the real effectiveness of moderation interventions remains unclear due to the high cost of data collection and…

人工智能 · 计算机科学 2026-04-01 Giacomo Fidone , Lucia Passaro , Riccardo Guidotti

Hate speech is one type of harmful online content which directly attacks or promotes hate towards a group or an individual member based on their actual or perceived aspects of identity, such as ethnicity, religion, and sexual orientation.…

计算与语言 · 计算机科学 2021-02-18 Wenjie Yin , Arkaitz Zubiaga

As AI systems become more integrated into daily life, the need for safer and more reliable moderation has never been greater. Large Language Models (LLMs) have demonstrated remarkable capabilities, surpassing earlier models in complexity…

人工智能 · 计算机科学 2026-01-13 Naseem Machlovi , Maryam Saleki , Innocent Ababio , Ruhul Amin

The automatic identification of harmful content online is of major concern for social media platforms, policymakers, and society. Researchers have studied textual, visual, and audio content, but typically in isolation. Yet, harmful content…

Pre-trained language models (PLMs) are trained on data that inherently contains gender biases, leading to undesirable impacts. Traditional debiasing methods often rely on external corpora, which may lack quality, diversity, or demographic…

计算与语言 · 计算机科学 2025-03-13 Liu Yu , Ludie Guo , Ping Kuang , Fan Zhou

As internet access expands, so does exposure to harmful content, increasing the need for effective moderation. Research has demonstrated that large language models (LLMs) can be effectively utilized for social media moderation tasks,…

计算与语言 · 计算机科学 2026-02-06 Hsuan-Yu Chou , Wajiha Naveed , Shuyan Zhou , Xiaowei Yang

We use structural topic modeling to examine racial bias in data collected to train models to detect hate speech and abusive language in social media posts. We augment the abusive language dataset by adding an additional feature indicating…

计算与语言 · 计算机科学 2020-05-28 Thomas Davidson , Debasmita Bhattacharya