中文
相关论文

相关论文: Detect Toxic Content to Improve Online Conversatio…

200 篇论文

Language models trained on large-scale unfiltered datasets curated from the open web acquire systemic biases, prejudices, and harmful views from their training data. We present a methodology for programmatically identifying and removing…

计算与语言 · 计算机科学 2021-11-30 Helen Ngo , Cooper Raterink , João G. M. Araújo , Ivan Zhang , Carol Chen , Adrien Morisot , Nicholas Frosst

Content moderation on social media platforms shapes the dynamics of online discourse, influencing whose voices are amplified and whose are suppressed. Recent studies have raised concerns about the fairness of content moderation practices,…

计算与语言 · 计算机科学 2024-06-25 Rebecca Dorn , Lee Kezar , Fred Morstatter , Kristina Lerman

Social media are an essential source of meaningful data that can be used in different tasks such as sentiment analysis and emotion recognition. Mostly, these tasks are solved with deep learning methods. Due to the fuzzy nature of textual…

计算与语言 · 计算机科学 2021-07-13 Olha Kaminska , Chris Cornelis , Veronique Hoste

In recent years, online social networks have allowed worldwide users to meet and discuss. As guarantors of these communities, the administrators of these platforms must prevent users from adopting inappropriate behaviors. This verification…

信息检索 · 计算机科学 2019-06-17 Noé Cecillon , Vincent Labatut , Richard Dufour , Georges Linarès

Online texts with toxic content are a clear threat to the users on social media in particular and society in general. Although many platforms have adopted various measures (e.g., machine learning-based hate-speech detection systems) to…

机器学习 · 计算机科学 2025-04-29 Yiran Ye , Thai Le , Dongwon Lee

This paper tackles one of the greatest limitations in Machine Learning: Data Scarcity. Specifically, we explore whether high accuracy classifiers can be built from small datasets, utilizing a combination of data augmentation techniques and…

计算与语言 · 计算机科学 2020-07-03 Chetanya Rastogi , Nikka Mofid , Fang-I Hsiao

The fact that everyone with a social media account can create and share content, and the increasing public reliance on social media platforms as a news and information source bring about significant challenges such as misinformation, fake…

计算与语言 · 计算机科学 2025-03-25 Muhidin A. Mohamed , Shuab D. Ahmed , Yahye A. Isse , Hanad M. Mohamed , Fuad M. Hassan , Houssein A. Assowe

Despite the extensive communication benefits offered by social media platforms, numerous challenges must be addressed to ensure user safety. One of the most significant risks faced by users on these platforms is targeted hate speech. Social…

计算与语言 · 计算机科学 2024-07-18 Sadar Jaf , Basel Barakat

Democratic societies need reliable information. Misinformation in popular media, such as news articles or videos, threatens to impair civic discourse. Citizens are, unfortunately, not equipped to verify the flood of content consumed daily…

计算与语言 · 计算机科学 2026-03-06 Tim Wittenborg , Constantin Sebastian Tremel , Markus Stocker , Sören Auer

We propose the Data Contamination Quiz (DCQ), a simple and effective approach to detect data contamination in large language models (LLMs) and estimate the amount of it. Specifically, we frame data contamination detection as a series of…

计算与语言 · 计算机科学 2025-04-29 Shahriar Golchin , Mihai Surdeanu

Detecting toxic language including sexism, harassment and abusive behaviour, remains a critical challenge, particularly in its subtle and context-dependent forms. Existing approaches largely focus on isolated message-level classification,…

The enormous use of sarcastic text in all forms of communication in social media will have a physiological effect on target users. Each user has a different approach to misusing and recognising sarcasm. Sarcasm detection is difficult even…

计算与语言 · 计算机科学 2023-04-19 Swapnil Mane , Vaibhav Khatavkar

The utilization of social media material in journalistic workflows is increasing, demanding automated methods for the identification of mis- and disinformation. Since textual contradiction across social media posts can be a signal of…

计算与语言 · 计算机科学 2017-07-12 Piroska Lendvai , Uwe D. Reichel

Social media datasets are essential for research on a variety of topics, such as disinformation, influence operations, hate speech detection, or influencer marketing practices. However, access to social media datasets is often constrained…

计算与语言 · 计算机科学 2025-05-07 Henry Tari , Nojus Sereiva , Rishabh Kaushal , Thales Bertaglia , Adriana Iamnitchi

Conspiracy theories erode public trust in science and institutions while resisting debunking by evolving and absorbing counter-evidence. As AI-generated misinformation becomes increasingly sophisticated, understanding rhetorical patterns in…

计算与语言 · 计算机科学 2025-08-29 Luke Bates , Max Glockner , Preslav Nakov , Iryna Gurevych

The remarkable ability of diffusion models to generate high-fidelity images has led to their widespread adoption. However, concerns have also arisen regarding their potential to produce Not Safe for Work (NSFW) content and exhibit social…

计算与语言 · 计算机科学 2025-05-22 Zhiwen Li , Die Chen , Mingyuan Fan , Cen Chen , Yaliang Li , Yanhao Wang , Wenmeng Zhou

Fake news, rumor, incorrect information, and misinformation detection are nowadays crucial issues as these might have serious consequences for our social fabrics. The rate of such information is increasing rapidly due to the availability of…

计算与语言 · 计算机科学 2018-11-13 Arjun Roy , Kingshuk Basak , Asif Ekbal , Pushpak Bhattacharyya

The ability of large language models to generate complex texts allows them to be widely integrated into many aspects of life, and their output can quickly fill all network resources. As the impact of LLMs grows, it becomes increasingly…

计算与语言 · 计算机科学 2024-11-12 Yongye Su , Yuqing Wu

Text-based misinformation permeates online discourses, yet evidence of people's ability to discern truth from such deceptive textual content is scarce. We analyze a novel TV game show data where conversations in a high-stake environment…

计算与语言 · 计算机科学 2024-04-09 Sanchaita Hazra , Bodhisattwa Prasad Majumder

Recent research in opinion mining proposed word embedding-based topic modeling methods that provide superior coherence compared to traditional topic modeling. In this paper, we demonstrate how these methods can be used to display correlated…