中文
相关论文

相关论文: Understanding Multi-Turn Toxic Behaviors in Open-D…

200 篇论文

Recent advances in large language models (LLMs) have led to the development of powerful AI chatbots capable of engaging in natural and human-like conversations. However, these chatbots can be potentially harmful, exhibiting manipulative,…

人工智能 · 计算机科学 2023-04-04 Baihan Lin , Djallel Bouneffouf , Guillermo Cecchi , Kush R. Varshney

Large language models (LLMs) have shown incredible capabilities and transcended the natural language processing (NLP) community, with adoption throughout many services like healthcare, therapy, education, and customer service. Since users…

计算与语言 · 计算机科学 2023-04-12 Ameet Deshpande , Vishvak Murahari , Tanmay Rajpurohit , Ashwin Kalyan , Karthik Narasimhan

Fostering a collaborative and inclusive environment is crucial for the sustained progress of open source development. However, the prevalence of negative discourse, often manifested as toxic comments, poses significant challenges to…

软件工程 · 计算机科学 2023-12-21 Shyamal Mishra , Preetha Chatterjee

Despite remarkable advances that large language models have achieved in chatbots, maintaining a non-toxic user-AI interactive environment has become increasingly critical nowadays. However, previous efforts in toxicity detection have been…

计算与语言 · 计算机科学 2023-10-27 Zi Lin , Zihan Wang , Yongqi Tong , Yangkun Wang , Yuxin Guo , Yujia Wang , Jingbo Shang

Dialogue models trained on human conversations inadvertently learn to generate toxic responses. In addition to producing explicitly offensive utterances, these models can also implicitly insult a group or individual by aligning themselves…

计算与语言 · 计算机科学 2021-09-14 Ashutosh Baheti , Maarten Sap , Alan Ritter , Mark Riedl

Chatbots are one class of intelligent, conversational software agents activated by natural language input (which can be in the form of text, voice, or both). They provide conversational output in response, and if commanded, can sometimes…

计算机与社会 · 计算机科学 2017-04-18 Nicole M. Radziwill , Morgan C. Benton

The detection of offensive language in the context of a dialogue has become an increasingly important application of natural language processing. The detection of trolls in public forums (Gal\'an-Garc\'ia et al., 2016), and the deployment…

计算与语言 · 计算机科学 2019-08-20 Emily Dinan , Samuel Humeau , Bharath Chintagunta , Jason Weston

Politically sensitive topics are still a challenge for open-domain chatbots. However, dealing with politically sensitive content in a responsible, non-partisan, and safe behavior way is integral for these chatbots. Currently, the main…

计算与语言 · 计算机科学 2021-06-14 Yejin Bang , Nayeon Lee , Etsuko Ishii , Andrea Madotto , Pascale Fung

Developing high-performing dialogue systems benefits from the automatic identification of undesirable behaviors in system responses. However, detecting such behaviors remains challenging, as it draws on a breadth of general knowledge and…

计算与语言 · 计算机科学 2023-09-14 Sarah E. Finch , Ellie S. Paek , Jinho D. Choi

Recent advances in interactive large language models like ChatGPT have revolutionized various domains; however, their behavior in natural and role-play conversation settings remains underexplored. In our study, we address this gap by deeply…

计算与语言 · 计算机科学 2024-03-28 Yufei Tao , Ameeta Agrawal , Judit Dombi , Tetyana Sydorenko , Jung In Lee

The growing deployment of large language model (LLM) based agents that interact with external environments has created new attack surfaces for adversarial manipulation. One major threat is indirect prompt injection, where attackers embed…

计算与语言 · 计算机科学 2026-04-14 Hwan Chang , Yonghyun Jun , Hwanhee Lee

Question-and-answer agents like ChatGPT offer a novel tool for use as a potential honeypot interface in cyber security. By imitating Linux, Mac, and Windows terminal commands and providing an interface for TeamViewer, nmap, and ping, it is…

密码学与安全 · 计算机科学 2023-01-11 Forrest McKee , David Noever

The emergence of Generative AI (Gen AI) and Large Language Models (LLMs) has enabled more advanced chatbots capable of human-like interactions. However, these conversational agents introduce a broader set of operational risks that extend…

密码学与安全 · 计算机科学 2025-05-09 Pedro Pinacho-Davidson , Fernando Gutierrez , Pablo Zapata , Rodolfo Vergara , Pablo Aqueveque

This study explores real-world human interactions with large language models (LLMs) in diverse, unconstrained settings in contrast to most prior research focusing on ethically trimmed models like ChatGPT for specific tasks. We aim to…

人机交互 · 计算机科学 2024-07-09 Johannes Schneider , Arianna Casanova Flores , Anne-Catherine Kranz

The study illustrates a first step towards an ongoing work aimed at developing a dataset of dialogues potentially useful for customer service conversation management between humans and AI chatbots. The approach exploits ChatGPT 3.5 to…

人机交互 · 计算机科学 2025-01-03 Alfredo Cuzzocrea , Giovanni Pilato , Pablo Garcia Bringas

In the rapidly evolving domain of artificial intelligence, chatbots have emerged as a potent tool for various applications ranging from e-commerce to healthcare. This research delves into the intricacies of chatbot technology, from its…

人机交互 · 计算机科学 2023-11-17 Feriel Khennouche , Youssef Elmir , Nabil Djebari , Yassine Himeur , Abbes Amira

Recent progress on neural approaches for language processing has triggered a resurgence of interest on building intelligent open-domain chatbots. However, even the state-of-the-art neural chatbots cannot produce satisfying responses for…

计算与语言 · 计算机科学 2022-08-10 Behnam Hedayatnia , Di Jin , Yang Liu , Dilek Hakkani-Tur

AI-driven chatbots such as ChatGPT have caused a tremendous hype lately. For BPM applications, several applications for AI-driven chatbots have been identified to be promising to generate business value, including explanation of process…

计算与语言 · 计算机科学 2024-01-19 Nataliia Klievtsova , Janik-Vasily Benzin , Timotheus Kampik , Juergen Mangler , Stefanie Rinderle-Ma

As mental health chatbots proliferate to address the global treatment gap, a critical question emerges: How do we design for relational safety the quality of interaction patterns that unfold across conversations rather than the correctness…

人机交互 · 计算机科学 2026-02-27 Joydeep Chandra , Satyam Kumar Navneet , Yong Zhang

This study evaluates the effectiveness of ChatGPT, an advanced AI model for natural language processing, in identifying targeting and inappropriate language in online comments. With the increasing challenge of moderating vast volumes of…

计算与语言 · 计算机科学 2025-05-29 Barbarestani Baran , Maks Isa , Vossen Piek