中文
相关论文

相关论文: MTTM: Metamorphic Testing for Textual Content Mode…

200 篇论文

The proliferation of harmful content on online platforms is a major societal problem, which comes in many different forms including hate speech, offensive language, bullying and harassment, misinformation, spam, violence, graphic content,…

Even with various regulations in place across countries and social media platforms (Government of India, 2021; European Parliament and Council of the European Union, 2022, digital abusive speech remains a significant issue. One potential…

In today's digital world, social media plays a significant role in facilitating communication and content sharing. However, the exponential rise in user-generated content has led to challenges in maintaining a respectful online environment.…

计算与语言 · 计算机科学 2024-03-05 Mohammad Dehghani

Social media platforms have been establishing content moderation guidelines and employing various moderation policies to counter hate speech and misinformation. The goal of this paper is to study these community guidelines and moderation…

计算机与社会 · 计算机科学 2023-08-10 Mohit Singhal , Chen Ling , Pujan Paudel , Poojitha Thota , Nihal Kumarswamy , Gianluca Stringhini , Shirin Nilizadeh

The proliferation of harmful content shared online poses a threat to online information integrity and the integrity of discussion across platforms. Despite various moderation interventions adopted by social media platforms, researchers and…

社会与信息网络 · 计算机科学 2023-04-07 Valerio La Gatta , Luca Luceri , Francesco Fabbri , Emilio Ferrara

Large language models (LLMs) are increasingly being used for emotional support. They are also being developed for formal therapy purposes. However, LLMs like ChaptGPT or Llama are often developed with content moderation guardrails that…

人机交互 · 计算机科学 2026-05-26 Jiwon Kim , Claire Wang , Taeung Yoon , Sabelle Huang , Koustuv Saha

Given the societal impact of unsafe content generated by large language models (LLMs), ensuring that LLM services comply with safety standards is a crucial concern for LLM service providers. Common content moderation methods are limited by…

计算与语言 · 计算机科学 2024-09-06 Jialin Wu , Jiangyi Deng , Shengyuan Pang , Yanjiao Chen , Jiayang Xu , Xinfeng Li , Wenyuan Xu

The proliferation of harmful and offensive content is a problem that many online platforms face today. One of the most common approaches for moderating offensive content online is via the identification and removal after it has been posted,…

社会与信息网络 · 计算机科学 2021-12-03 Matthew Katsaros , Kathy Yang , Lauren Fratamico

Offensive Language detection in social media platforms has been an active field of research over the past years. In non-native English spoken countries, social media users mostly use a code-mixed form of text in their posts/comments. This…

计算与语言 · 计算机科学 2022-12-13 Charangan Vasantharajan , Uthayasanker Thayasivam

To protect users from massive hateful content, existing works studied automated hate speech detection. Despite the existing efforts, one question remains: do automated hate speech detectors conform to social media content policies? A…

Polarization, declining trust, and wavering support for democratic norms are pressing threats to U.S. democracy. Exposure to verified and quality news may lower individual susceptibility to these threats and make citizens more resilient to…

社会与信息网络 · 计算机科学 2024-04-02 Hadi Askari , Anshuman Chhabra , Bernhard Clemm von Hohenberg , Michael Heseltine , Magdalena Wojcieszak

This study evaluates the effectiveness of ChatGPT, an advanced AI model for natural language processing, in identifying targeting and inappropriate language in online comments. With the increasing challenge of moderating vast volumes of…

计算与语言 · 计算机科学 2025-05-29 Barbarestani Baran , Maks Isa , Vossen Piek

Social media platforms have become critical spaces for discussing mental health concerns, including eating disorders. While these platforms can provide valuable support networks, they may also amplify harmful content that glorifies…

社会与信息网络 · 计算机科学 2024-12-23 Kristina Lerman , Minh Duc Chu , Charles Bickham , Luca Luceri , Emilio Ferrara

Online conversations can be toxic and subjected to threats, abuse, or harassment. To identify toxic text comments, several deep learning and machine learning models have been proposed throughout the years. However, recent studies…

机器学习 · 计算机科学 2023-11-09 Md Azim Khan

Detecting online sexual predatory behaviours and abusive language on social media platforms has become a critical area of research due to the growing concerns about online safety, especially for vulnerable populations such as children and…

计算与语言 · 计算机科学 2023-08-29 Thanh Thi Nguyen , Campbell Wilson , Janis Dalins

This study explores real-world human interactions with large language models (LLMs) in diverse, unconstrained settings in contrast to most prior research focusing on ethically trimmed models like ChatGPT for specific tasks. We aim to…

人机交互 · 计算机科学 2024-07-09 Johannes Schneider , Arianna Casanova Flores , Anne-Catherine Kranz

As the volume of video content online grows exponentially, the demand for moderation of unsafe videos has surpassed human capabilities, posing both operational and mental health challenges. While recent studies demonstrated the merits of…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Adi Levi , Or Levi , Sardhendu Mishra , Jonathan Morra

Unlike regular tokens derived from existing text corpora, special tokens are artificially created to annotate structured conversations during the fine-tuning process of Large Language Models (LLMs). Serving as metadata of training data,…

密码学与安全 · 计算机科学 2025-10-14 Wentian Zhu , Zhen Xiang , Wei Niu , Le Guan

Large language models (LLMs) have introduced substantial challenges to software quality assurance due to their generative, probabilistic, and open-ended nature, which intensifies the oracle problem and limits the applicability of…

软件工程 · 计算机科学 2026-05-15 Zheng Zheng , Zenghui Zhou , Yinwang Xu , Daixu Ren , Tsong Yueh Chen

Hate Speech has become a major content moderation issue for online social media platforms. Given the volume and velocity of online content production, it is impossible to manually moderate hate speech related content on any platform. In…

计算与语言 · 计算机科学 2021-01-28 Sudhanshu Mishra , Shivangi Prasad , Shubhanshu Mishra