中文
相关论文

相关论文: Mitigating Toxic Degeneration with Empathetic Data…

200 篇论文

Recent developments in generative AI have shone a spotlight on high-performance synthetic text generation technologies. The now wide availability and ease of use of such models highlights the urgent need to provide equally powerful…

计算与语言 · 计算机科学 2023-10-25 Alan Cowap , Yvette Graham , Jennifer Foster

The collection and curation of high-quality training data is crucial for developing text classification models with superior performance, but it is often associated with significant costs and time investment. Researchers have recently…

计算与语言 · 计算机科学 2023-10-16 Zhuoyan Li , Hangxiao Zhu , Zhuoran Lu , Ming Yin

In empathetic conversations, humans express their empathy to others with empathetic intents. However, most existing empathetic conversational methods suffer from a lack of empathetic intents, which leads to monotonous empathy. To address…

计算与语言 · 计算机科学 2022-04-27 Mao Yan Chen , Siheng Li , Yujiu Yang

Detection of some types of toxic language is hampered by extreme scarcity of labeled training data. Data augmentation - generating new synthetic data from a labeled seed dataset - can help. The efficacy of data augmentation on toxic…

计算与语言 · 计算机科学 2020-10-27 Mika Juuti , Tommi Gröndahl , Adrian Flanagan , N. Asokan

Empathetic dialogue is a human-like behavior that requires the perception of both affective factors (e.g., emotion status) and cognitive factors (e.g., cause of the emotion). Besides concerning emotion status in early work, the latest…

计算与语言 · 计算机科学 2023-02-24 Yushan Qian , Bo Wang , Ting-En Lin , Yinhe Zheng , Ying Zhu , Dongming Zhao , Yuexian Hou , Yuchuan Wu , Yongbin Li

The rise of online communication platforms has been accompanied by some undesirable effects, such as the proliferation of aggressive and abusive behaviour online. Aiming to tackle this problem, the natural language processing (NLP)…

计算与语言 · 计算机科学 2020-05-29 Santhosh Rajamanickam , Pushkar Mishra , Helen Yannakoudakis , Ekaterina Shutova

Moderating offensive, hateful, and toxic language has always been an important but challenging topic in the domain of safe use in NLP. The emerging large language models (LLMs), such as ChatGPT, can potentially further accentuate this…

计算机与社会 · 计算机科学 2023-11-28 Boyang Zhang , Xinyue Shen , Wai Man Si , Zeyang Sha , Zeyuan Chen , Ahmed Salem , Yun Shen , Michael Backes , Yang Zhang

Increasingly larger datasets have become a standard ingredient to advancing the state-of-the-art in NLP. However, data quality might have already become the bottleneck to unlock further gains. Given the diversity and the sizes of modern…

计算与语言 · 计算机科学 2023-10-18 Irina Bejan , Artem Sokolov , Katja Filippova

Advancements in emotion aware language processing increasingly shape vital NLP applications ranging from conversational AI and affective computing to computational psychology and creative content generation. Existing emotion datasets either…

计算与语言 · 计算机科学 2025-04-14 Vishal Gandhi , Sagar Gandhi

Language models are the new state-of-the-art natural language processing (NLP) models and they are being increasingly used in many NLP tasks. Even though there is evidence that language models are biased, the impact of that bias on the…

计算与语言 · 计算机科学 2024-04-29 Fatma Elsafoury , Stamos Katsigiannis

Emotion detection is a central problem in NLP, with recent progress driven by transformer-based models trained on established datasets. However, little is known about the linguistic regularities that characterize how emotions are expressed…

计算与语言 · 计算机科学 2026-03-24 Florian Lecourt , Madalina Croitoru , Konstantin Todorov

Augmenting toxic language data in a controllable and class-specific manner is crucial for improving robustness in toxicity classification, yet remains challenging due to limited supervision and distributional skew. We propose ToxiGAN, a…

计算与语言 · 计算机科学 2026-01-07 Peiran Li , Jan Fillies , Adrian Paschke

There are a number of diverging hypotheses about the neural text degeneration problem, i.e., generating repetitive and dull loops, which makes this problem both interesting and confusing. In this work, we aim to advance our understanding by…

计算与语言 · 计算机科学 2023-10-17 Huayang Li , Tian Lan , Zihao Fu , Deng Cai , Lemao Liu , Nigel Collier , Taro Watanabe , Yixuan Su

This paper delves into enhancing the classification performance on the GoEmotions dataset, a large, manually annotated dataset for emotion detection in text. The primary goal of this paper is to address the challenges of detecting subtle…

计算与语言 · 计算机科学 2024-04-10 Kaipeng Wang , Zhi Jing , Yongye Su , Yikun Han

Mental manipulation, a significant form of abuse in interpersonal conversations, presents a challenge to identify due to its context-dependent and often subtle nature. The detection of manipulative language is essential for protecting…

计算与语言 · 计算机科学 2024-05-28 Yuxin Wang , Ivory Yang , Saeed Hassanpour , Soroush Vosoughi

Text data can pose a risk of harm. However, the risks are not fully understood, and how to handle, present, and discuss harmful text in a safe way remains an unresolved issue in the NLP community. We provide an analytical framework…

计算与语言 · 计算机科学 2023-02-28 Hannah Rose Kirk , Abeba Birhane , Bertie Vidgen , Leon Derczynski

Large language models (LLMs) rely on web-scale corpora for pre-training. The noise inherent in these datasets tends to obscure meaningful patterns and ultimately degrade model performance. Data curation mitigates but cannot eliminate such…

计算与语言 · 计算机科学 2026-05-12 Xu Guo , Runyu Peng , Jian Tong , Yunhua Zhou , Haijun Lv , Zhihui Lu , Qipeng Guo

This manuscript presents a methodical examination of the utilization of Artificial Intelligence in the assessment of emotions in texts related to healthcare, with a particular focus on the incorporation of Natural Language Processing and…

计算与语言 · 计算机科学 2024-04-22 Prashant Kumar Nag , Amit Bhagat , R. Vishnu Priya , Deepak kumar Khare

Large language models (LLMs) have become integral to our professional workflows and daily lives. Nevertheless, these machine companions of ours have a critical flaw: the huge amount of data which endows them with vast and diverse knowledge,…

计算与语言 · 计算机科学 2024-05-21 Tinh Son Luong , Thanh-Thien Le , Linh Ngo Van , Thien Huu Nguyen

Large pre-trained language models are often trained on large volumes of internet data, some of which may contain toxic or abusive language. Consequently, language models encode toxic information, which makes the real-world usage of these…

计算与语言 · 计算机科学 2021-12-16 Andrew Wang , Mohit Sudhakar , Yangfeng Ji