中文
相关论文

相关论文: Toxicity Detection Should Measure Contextual Harm,…

200 篇论文

Detecting toxicity in multimodal data remains a significant challenge, as harmful meanings often lurk beneath seemingly benign individual modalities: only emerging when modalities are combined and semantic associations are activated. To…

机器学习 · 计算机科学 2026-02-04 Guanzong Wu , Zihao Zhu , Siwei Lyu , Baoyuan Wu

Toxic language detection is crucial for creating safer online environments and limiting the spread of harmful content. While toxic language detection has been under-explored in Persian, the current work compares different methods for this…

计算与语言 · 计算机科学 2025-06-05 Zahra Bokaei , Walid Magdy , Bonnie Webber

We present Contamination Detection via Context (CoDeC), a practical and accurate method to detect and quantify training data contamination in large language models. CoDeC distinguishes between data memorized during training and data outside…

计算与语言 · 计算机科学 2026-05-13 Michał Zawalski , Meriem Boubdir , Klaudia Bałazy , Besmira Nushi , Pablo Ribalta

Given the dynamic nature of toxic language use, automated methods for detecting toxic spans are likely to encounter distributional shift. To explore this phenomenon, we evaluate three approaches for detecting toxic spans under cross-domain…

计算与语言 · 计算机科学 2023-06-19 Stefan F. Schouten , Baran Barbarestani , Wondimagegnhue Tufa , Piek Vossen , Ilia Markov

Existing toxic detection models face significant limitations, such as lack of transparency, customization, and reproducibility. These challenges stem from the closed-source nature of their training data and the paucity of explanations for…

计算与语言 · 计算机科学 2025-01-24 Tinh Son Luong , Thanh-Thien Le , Thang Viet Doan , Linh Ngo Van , Thien Huu Nguyen , Diep Thi-Ngoc Nguyen

New text as data techniques offer a great promise: the ability to inductively discover measures that are useful for testing social science theories of interest from large collections of text. We introduce a conceptual framework for making…

This work proposes a contextualised detection framework for implicitly hateful speech, implemented as a multi-agent system comprising a central Moderator Agent and dynamically constructed Community Agents representing specific demographic…

计算与语言 · 计算机科学 2026-01-28 Ewelina Gajewska , Katarzyna Budzynska , Jarosław A Chudziak

As large language models become integral to high-stakes applications, ensuring their robustness and fairness is critical. Despite their success, large language models remain vulnerable to adversarial attacks, where small perturbations, such…

人工智能 · 计算机科学 2026-02-02 Danqing Chen , Tobias Ladner , Ahmed Rayen Mhadhbi , Matthias Althoff

Toxic speech, also known as hate speech, is regarded as one of the crucial issues plaguing online social media today. Most recent work on toxic speech detection is constrained to the modality of text and written conversations with very…

计算与语言 · 计算机科学 2022-04-05 Sreyan Ghosh , Samden Lepcha , S Sakshi , Rajiv Ratn Shah , S. Umesh

Social media platforms have evolved rapidly in modernity without strong regulation. One clear obstacle faced by current users is that of toxicity. Toxicity on social media manifests through a number of forms, including harassment,…

社会与信息网络 · 计算机科学 2024-10-30 Rhett Hanscom , Tamara Silbergleit Lehman , Qin Lv , Shivakant Mishra

Text classification is an important topic in the field of natural language processing. It has been preliminarily applied in information retrieval, digital library, automatic abstracting, text filtering, word semantic discrimination and many…

计算与语言 · 计算机科学 2023-12-20 Hao Li , Brandon Bennett

We present our works on SemEval-2021 Task 5 about Toxic Spans Detection. This task aims to build a model for identifying toxic words in whole posts. We use the BiLSTM-CRF model combining with ToxicBERT Classification to train the detection…

计算与语言 · 计算机科学 2021-08-02 Son T. Luu , Ngan Luu-Thuy Nguyen

Hate speech detection refers to the task of detecting hateful content that aims at denigrating an individual or a group based on their religion, gender, sexual orientation, or other characteristics. Due to the different policies of the…

计算与语言 · 计算机科学 2023-10-10 Paras Sheth , Tharindu Kumarage , Raha Moraffah , Aman Chadha , Huan Liu

Harmful text detection has become a crucial task in the development and deployment of large language models, especially as AI-generated content continues to expand across digital platforms. This study proposes a joint retrieval framework…

计算与语言 · 计算机科学 2025-04-04 Zidong Yu , Shuo Wang , Nan Jiang , Weiqiang Huang , Xu Han , Junliang Du

Identifying and understanding quality phrases from context is a fundamental task in text mining. The most challenging part of this task arguably lies in uncommon, emerging, and domain-specific phrases. The infrequent nature of these phrases…

计算与语言 · 计算机科学 2021-06-01 Xiaotao Gu , Zihan Wang , Zhenyu Bi , Yu Meng , Liyuan Liu , Jiawei Han , Jingbo Shang

Toxic conversations during software development interactions may have serious repercussions on a Free and Open Source Software (FOSS) development project. For example, victims of toxic conversations may become afraid to express themselves,…

软件工程 · 计算机科学 2023-02-09 Jaydeb Sarker , Asif Kamal Turzo , Ming Dong , Amiangshu Bosu

To support safety and inclusion in online communications, significant efforts in NLP research have been put towards addressing the problem of abusive content detection, commonly defined as a supervised classification task. The research…

计算与语言 · 计算机科学 2020-10-29 Svetlana Kiritchenko , Isar Nejadgholi

Lack of moderation in online communities enables participants to incur in personal aggression, harassment or cyberbullying, issues that have been accentuated by extremist radicalisation in the contemporary post-truth politics scenario. This…

计算与语言 · 计算机科学 2018-01-08 Nestor Rodriguez , Sergio Rojas-Galeano

Large language models often face a three-way trade-off among detection accuracy, inference latency, and deployment cost when used in real-world safety-sensitive applications. This paper introduces Prefix Probing, a black-box harmful content…

人工智能 · 计算机科学 2025-12-19 Jirui Yang , Hengqi Guo , Zhihui Lu , Yi Zhao , Yuansen Zhang , Shijing Hu , Qiang Duan , Yinggui Wang , Tao Wei

The abstract outlines the problem of toxic comments on social media platforms, where individuals use disrespectful, abusive, and unreasonable language that can drive users away from discussions. This behavior is referred to as anti-social…

机器学习 · 计算机科学 2023-04-17 K. Poojitha , A. Sai Charish , M. Arun Kuamr Reddy , S. Ayyasamy