中文
相关论文

相关论文: Towards Detecting Contextual Real-Time Toxicity fo…

200 篇论文

We present Moderator, a policy-based model management system that allows administrators to specify fine-grained content moderation policies and modify the weights of a text-to-image (TTI) model to make it significantly more challenging for…

密码学与安全 · 计算机科学 2024-09-13 Peiran Wang , Qiyu Li , Longxuan Yu , Ziyao Wang , Ang Li , Haojian Jin

This paper envisions a multi-agent system for detecting the presence of hate speech in online social media platforms such as Twitter and Facebook. We introduce a novel framework employing deep learning techniques to coordinate the channels…

人工智能 · 计算机科学 2021-05-05 Gaurav Sahu , Robin Cohen , Olga Vechtomova

In the wake of a polarizing election, the cyber world is laden with hate speech. Context accompanying a hate speech text is useful for identifying hate speech, which however has been largely overlooked in existing datasets and hate speech…

计算与语言 · 计算机科学 2018-05-23 Lei Gao , Ruihong Huang

Internet memes are a powerful form of online communication, yet their nature and reliance on commonsense knowledge make toxicity detection challenging. Identifying key features for meme interpretation and understanding, is a crucial task.…

计算与语言 · 计算机科学 2026-03-05 Stefano De Giorgis , Ting-Chih Chen , Filip Ilievski

The exponential growth of social media platforms has brought about a revolution in communication and content dissemination in human society. Nevertheless, these platforms are being increasingly misused to spread toxic content, including…

软件工程 · 计算机科学 2023-08-22 Wenxuan Wang , Jingyuan Huang , Jen-tse Huang , Chang Chen , Jiazhen Gu , Pinjia He , Michael R. Lyu

Toxicity on GitHub can severely impact Open Source Software (OSS) development communities. To mitigate such behavior, a better understanding of its nature and how various measurable characteristics of project contexts and participants are…

软件工程 · 计算机科学 2025-02-13 Jaydeb Sarker , Asif Kamal Turzo , Amiangshu Bosu

Machine-generated text (MGT) detection is critical for regulating online information ecosystems, yet existing detectors often underperform in few-shot settings and remain vulnerable to adversarial, humanizing attacks. To build accurate and…

密码学与安全 · 计算机科学 2026-05-05 Wenjing Duan , Qi Zhou , Yuanfan Li

Interactive systems such as chatbots and games are increasingly used to persuade and educate on sustainability-related topics, yet it remains unclear how different delivery formats shape learning and persuasive outcomes when content is held…

人机交互 · 计算机科学 2026-02-25 Seyed Hossein Alavi , Zining Wang , Shruthi Chockkalingam , Raymond T. Ng , Vered Shwartz

We present the Multi-Modal Discussion Transformer (mDT), a novel methodfor detecting hate speech in online social networks such as Reddit discussions. In contrast to traditional comment-only methods, our approach to labelling a comment as…

计算与语言 · 计算机科学 2024-02-23 Liam Hebert , Gaurav Sahu , Yuxuan Guo , Nanda Kishore Sreenivas , Lukasz Golab , Robin Cohen

The detection of offensive language in the context of a dialogue has become an increasingly important application of natural language processing. The detection of trolls in public forums (Gal\'an-Garc\'ia et al., 2016), and the deployment…

计算与语言 · 计算机科学 2019-08-20 Emily Dinan , Samuel Humeau , Bharath Chintagunta , Jason Weston

This paper proposes a new DeepFake detector FakeBuster for detecting impostors during video conferencing and manipulated faces on social media. FakeBuster is a standalone deep learning based solution, which enables a user to detect if…

计算机视觉与模式识别 · 计算机科学 2021-01-12 Vineet Mehta , Parul Gupta , Ramanathan Subramanian , Abhinav Dhall

Text-to-image (T2I) models such as Stable Diffusion have advanced rapidly and are now widely used in content creation. However, these models can be misused to generate harmful content, including nudity or violence, posing significant safety…

密码学与安全 · 计算机科学 2025-06-13 Zilong Wang , Xiang Zheng , Xiaosen Wang , Bo Wang , Xingjun Ma , Yu-Gang Jiang

The convenience of social media has also enabled its misuse, potentially resulting in toxic behavior. Nearly 66% of internet users have observed online harassment, and 41% claim personal experience, with 18% facing severe forms of online…

The spread of information through social media platforms can create environments possibly hostile to vulnerable communities and silence certain groups in society. To mitigate such instances, several models have been developed to detect hate…

To meet the demands of content moderation, online platforms have resorted to automated systems. Newer forms of real-time engagement($\textit{e.g.}$, users commenting on live streams) on platforms like Twitch exert additional pressures on…

计算与语言 · 计算机科学 2025-06-11 Prarabdh Shukla , Wei Yin Chong , Yash Patel , Brennan Schaffner , Danish Pruthi , Arjun Bhagoji

Although automated harmful content detection systems are frequently used to monitor online platforms, moderators and end users frequently cannot understand the logic underlying their predictions. While recent studies have focused on…

计算与语言 · 计算机科学 2026-03-20 Trishita Dhara , Siddhesh Sheth

Social media platforms have evolved rapidly in modernity without strong regulation. One clear obstacle faced by current users is that of toxicity. Toxicity on social media manifests through a number of forms, including harassment,…

社会与信息网络 · 计算机科学 2024-10-30 Rhett Hanscom , Tamara Silbergleit Lehman , Qin Lv , Shivakant Mishra

Conventional methods of assessing attitudes towards climate change are limited in capturing authentic opinions, primarily stemming from a lack of context-specific assessment strategies and an overreliance on simplistic surveys. Game-based…

人机交互 · 计算机科学 2024-03-28 Suifang Zhou , Latisha Besariani Hendra , Qinshi Zhang , Jussi Holopainen , RAY LC

This work proposes a contextualised detection framework for implicitly hateful speech, implemented as a multi-agent system comprising a central Moderator Agent and dynamically constructed Community Agents representing specific demographic…

计算与语言 · 计算机科学 2026-01-28 Ewelina Gajewska , Katarzyna Budzynska , Jarosław A Chudziak

The volume of machine-generated content online has grown dramatically due to the widespread use of Large Language Models (LLMs), leading to new challenges for content moderation systems. Conventional content moderation classifiers, which…

计算与语言 · 计算机科学 2026-05-26 Shaz Furniturewala , Arkaitz Zubiaga
‹ 上一页 1 8 9 10 下一页 ›