中文
相关论文

相关论文: Down the Toxicity Rabbit Hole: A Novel Framework t…

200 篇论文

While large language models (LLMs) present significant potential for supporting numerous real-world applications and delivering positive social impacts, they still face significant challenges in terms of the inherent risk of privacy…

人工智能 · 计算机科学 2025-01-17 Huandong Wang , Wenjie Fu , Yingzhou Tang , Zhilong Chen , Yuxi Huang , Jinghua Piao , Chen Gao , Fengli Xu , Tao Jiang , Yong Li

The proliferation of social media platforms has led to an increase in the spread of hate speech, particularly targeting vulnerable communities. Unfortunately, existing methods for automatically identifying and blocking toxic language rely…

计算与语言 · 计算机科学 2025-02-24 Shiza Ali , Jeremy Blackburn , Gianluca Stringhini

Large language models frequently generate toxic, hateful, or harmful content, yet existing mitigation methods rely on costly retraining or output-level filtering with no mechanistic insight into where toxicity originates internally. We…

计算与语言 · 计算机科学 2026-05-28 Himanshu Beniwal , Mayank Singh

Online abuse has grown increasingly complex, spanning toxic language, harassment, manipulation, and fraudulent behavior. Traditional machine-learning approaches dependent on static classifiers and labor-intensive labeling struggle to keep…

计算与语言 · 计算机科学 2026-04-02 Suraj Kath , Sanket Badhe , Preet Shah , Ashwin Sampathkumar , Shivani Gupta

Large Language Model (LLM) agents increasingly act through external tools, making their safety contingent on tool-call workflows rather than text generation alone. While recent benchmarks evaluate agents across diverse environments and risk…

软件工程 · 计算机科学 2026-03-20 Xuan Chen , Lu Yan , Ruqi Zhang , Xiangyu Zhang

The perceived toxicity of language can vary based on someone's identity and beliefs, but this variation is often ignored when collecting toxic language datasets, resulting in dataset and model biases. We seek to understand the who, why, and…

计算与语言 · 计算机科学 2022-05-11 Maarten Sap , Swabha Swayamdipta , Laura Vianna , Xuhui Zhou , Yejin Choi , Noah A. Smith

Large language models (LLMs) can exhibit concept-conditioned semantic divergence: common high-level cues (e.g., ideologies, public figures) elicit unusually uniform, stance-like responses that evade token-trigger audits. This behavior falls…

计算与语言 · 计算机科学 2026-01-28 Nay Myat Min , Long H. Pham , Yige Li , Jun Sun

Textual backdoor attacks are a kind of practical threat to NLP systems. By injecting a backdoor in the training phase, the adversary could control model predictions via predefined triggers. As various attack and defense models have been…

机器学习 · 计算机科学 2022-11-02 Ganqu Cui , Lifan Yuan , Bingxiang He , Yangyi Chen , Zhiyuan Liu , Maosong Sun

Toxic speech, also known as hate speech, is regarded as one of the crucial issues plaguing online social media today. Most recent work on toxic speech detection is constrained to the modality of text and written conversations with very…

计算与语言 · 计算机科学 2022-04-05 Sreyan Ghosh , Samden Lepcha , S Sakshi , Rajiv Ratn Shah , S. Umesh

Large Language Models (LLMs) such as ChatGPT, have gained significant attention due to their impressive natural language processing capabilities. It is crucial to prioritize human-centered principles when utilizing these models.…

计算与语言 · 计算机科学 2023-06-21 Yue Huang , Qihui Zhang , Philip S. Y , Lichao Sun

Large language models (LLMs) are increasingly used to make sense of ambiguous, open-textured, value-laden terms. Platforms routinely rely on LLMs for content moderation, asking them to label text based on disputed concepts like "hate…

计算机与社会 · 计算机科学 2026-03-09 Shira Gur-Arieh , Angelina Wang , Sina Fazelpour

The widespread use of Large Multimodal Models (LMMs) has raised concerns about model toxicity. However, current research mainly focuses on explicit toxicity, with less attention to some more implicit toxicity regarding prejudice and…

计算与语言 · 计算机科学 2025-05-26 Bohan Jin , Shuhan Qi , Kehai Chen , Xinyi Guo , Xuan Wang

As language models grow in popularity, it becomes increasingly important to clearly measure all possible markers of demographic identity in order to avoid perpetuating existing societal harms. Many datasets for measuring bias currently…

计算与语言 · 计算机科学 2022-10-31 Eric Michael Smith , Melissa Hall , Melanie Kambadur , Eleonora Presani , Adina Williams

Risk assessment tools are widely used around the country to inform decision making within the criminal justice system. Recently, considerable attention has been devoted to the question of whether such tools may suffer from racial bias. In…

统计方法学 · 统计学 2020-04-01 Riccardo Fogliato , Max G'Sell , Alexandra Chouldechova

Hate speech detection is a crucial area of research in natural language processing, essential for ensuring online community safety. However, detecting implicit hate speech, where harmful intent is conveyed in subtle or indirect ways,…

计算与语言 · 计算机科学 2025-04-17 Yumin Kim , Hwanhee Lee

Multilingual studies of social bias in open-ended LLM generation remain limited: most existing benchmarks are English-centric, template-based, or restricted to recognizing pre-specified stereotypes. We introduce StereoTales, a multilingual…

Large language models (LLMs) represent a major advance in artificial intelligence (AI) research. However, the widespread use of LLMs is also coupled with significant ethical and social challenges. Previous research has pointed towards…

计算与语言 · 计算机科学 2023-06-28 Jakob Mökander , Jonas Schuett , Hannah Rose Kirk , Luciano Floridi

Models trained on large unlabeled corpora of human interactions will learn patterns and mimic behaviors therein, which include offensive or otherwise toxic behavior and unwanted biases. We investigate a variety of methods to mitigate these…

计算与语言 · 计算机科学 2021-08-06 Jing Xu , Da Ju , Margaret Li , Y-Lan Boureau , Jason Weston , Emily Dinan

Large language models (LLMs) increasingly operate in high-stakes settings including healthcare and medicine, where demographic attributes such as race and ethnicity may be explicitly stated or implicitly inferred from text. However,…

计算与语言 · 计算机科学 2026-01-21 Shiyue Hu , Ruizhe Li , Yanjun Gao

Researchers and developers increasingly rely on toxicity scoring to moderate generative language model outputs, in settings such as customer service, information retrieval, and content generation. However, toxicity scoring may render…

人机交互 · 计算机科学 2024-04-23 Jennifer Chien , Kevin R. McKee , Jackie Kay , William Isaac