中文
相关论文

相关论文: Harmful Speech Detection by Language Models Exhibi…

200 篇论文

Social media communication has become a significant part of daily activity in modern societies. For this reason, ensuring safety in social media platforms is a necessity. Use of dangerous language such as physical threats in online…

计算与语言 · 计算机科学 2020-05-15 Ali Alshehri , El Moatez Billah Nagoudi , Muhammad Abdul-Mageed

Cyberbullying is a pervasive problem in online communities. To identify cyberbullying cases in large-scale social networks, content moderators depend on machine learning classifiers for automatic cyberbullying detection. However, existing…

社会与信息网络 · 计算机科学 2020-04-07 Caleb Ziems , Ymir Vigfusson , Fred Morstatter

Online marketplaces, while revolutionizing global commerce, have inadvertently facilitated the proliferation of illicit activities, including drug trafficking, counterfeit sales, and cybercrimes. Traditional content moderation methods such…

计算与语言 · 计算机科学 2026-03-06 Quoc Khoa Tran , Thanh Thi Nguyen , Campbell Wilson

Content moderation is a central mechanism through which platforms attempt to balance user engagement with community governance. Yet existing research has largely treated moderation as a uniform intervention, overlooking how moderator…

计算机与社会 · 计算机科学 2026-05-18 Siyi Zhou , Lindsay Young , Marlon Twyman , Emilio Ferrara

Nowadays, billions of people engage in communication and express their opinions on the internet daily. Unfortunately, not all of these expressions are friendly or compliant, making content moderation an indispensable task. A common approach…

机器学习 · 计算机科学 2024-03-08 Huan Ma , Changqing Zhang , Huazhu Fu , Peilin Zhao , Bingzhe Wu

The widespread application of Large Language Models (LLMs) involves ethical risks for users and societies. A prominent ethical risk of LLMs is the generation of unfair language output that reinforces or exacerbates harm for members of…

计算与语言 · 计算机科学 2025-03-03 Luise Mehner , Lena Alicija Philine Fiedler , Sabine Ammon , Dorothea Kolossa

This paper investigates the use of machine learning models for the classification of unhealthy online conversations containing one or more forms of subtler abuse, such as hostility, sarcasm, and generalization. We leveraged a public dataset…

计算与语言 · 计算机科学 2022-01-28 Shlok Gilda , Mirela Silva , Luiz Giovanini , Daniela Oliveira

Cultural context profoundly shapes how people interpret online content, yet vision-language models (VLMs) remain predominantly trained through Western or English-centric lenses. This limits their fairness and cross-cultural robustness in…

计算与语言 · 计算机科学 2026-02-13 Mo Wang , Kaixuan Ren , Pratik Jalan , Ahmed Ashraf , Tuong Vy Vu , Rahul Seetharaman , Shah Nawaz , Usman Naseem

To address the global challenge of online hate speech, prior research has developed detection models to flag such content on social media. However, due to systematic biases in evaluation datasets, the real-world effectiveness of these…

The rapid development of large language models (LLMs) has significantly improved the generation of fluent and convincing text, raising concerns about their potential misuse on social media platforms. We present a comprehensive methodology…

计算与语言 · 计算机科学 2025-01-22 Bryan E. Tuck , Rakesh M. Verma

The pervasive spread of misinformation and disinformation poses a significant threat to society. Professional fact-checkers play a key role in addressing this threat, but the vast scale of the problem forces them to prioritize their limited…

计算机与社会 · 计算机科学 2024-01-31 Terrence Neumann , Sooyong Lee , Maria De-Arteaga , Sina Fazelpour , Matthew Lease

Sexism in online content is a pervasive issue that necessitates effective classification techniques to mitigate its harmful impact. Online platforms often have sexist comments and posts that create a hostile environment, especially for…

计算与语言 · 计算机科学 2025-01-07 AmirMohammad Azadi , Baktash Ansari , Sina Zamani , Sauleh Eetemadi

This research offers a unique evaluation of how AI systems interpret the digital language of Generation Alpha (Gen Alpha, born 2010-2024). As the first cohort raised alongside AI, Gen Alpha faces new forms of online risk due to immersive…

计算机与社会 · 计算机科学 2025-05-19 Manisha Mehta , Fausto Giunchiglia

With the rise of social media, people can now form relationships and communities easily regardless of location, race, ethnicity, or gender. However, the power of social media simultaneously enables harmful online behavior such as harassment…

社会与信息网络 · 计算机科学 2016-06-28 Elaheh Raisi , Bert Huang

Large language models (LLMs) offer promising opportunities for organizational research. However, their built-in moderation systems can create problems when researchers try to analyze harmful content, often refusing to follow certain…

人工智能 · 计算机科学 2025-06-23 Mustafa Akben , Aaron Satko

Hate, derogatory, and offensive speech remains a persistent challenge in online platforms and public discourse. While automated detection systems are widely used, most focus on censorship or removal, raising concerns for transparency and…

The increasing accessibility of the internet facilitated social media usage and encouraged individuals to express their opinions liberally. Nevertheless, it also creates a place for content polluters to disseminate offensive posts or…

计算与语言 · 计算机科学 2021-03-02 Omar Sharif , Eftekhar Hossain , Mohammed Moshiul Hoque

Hate speech is a form of online harassment that involves the use of abusive language, and it is commonly seen in social media posts. This sort of harassment mainly focuses on specific group characteristics such as religion, gender,…

计算与语言 · 计算机科学 2022-06-10 Georgios K. Pitsilis

Online platforms face the challenge of moderating an ever-increasing volume of content, including harmful hate speech. In the absence of clear legal definitions and a lack of transparency regarding the role of algorithms in shaping…

计算机与社会 · 计算机科学 2024-06-21 David Hartmann , Amin Oueslati , Dimitri Staufer

This study evaluates the biases in Gemini 2.0 Flash Experimental, a state-of-the-art large language model (LLM) developed by Google, focusing on content moderation and gender disparities. By comparing its performance to ChatGPT-4o, examined…

计算与语言 · 计算机科学 2025-03-24 Roberto Balestri