English
Related papers

Related papers: MindGuard: Guardrail Classifiers for Multi-Turn Me…

200 papers

The development of conversational agents to interact with patients and deliver clinical advice has attracted the interest of many researchers, particularly in light of the COVID-19 pandemic. The training of an end-to-end neural based dialog…

Computation and Language · Computer Science 2022-12-13 Deeksha Varshney , Aizan Zafar , Niranshu Kumar Behra , Asif Ekbal

Malicious content generated by large language models (LLMs) can pose varying degrees of harm. Although existing LLM-based moderators can detect harmful content, they struggle to assess risk levels and may miss lower-risk outputs. Accurate…

Computation and Language · Computer Science 2025-03-11 Fan Yin , Philippe Laban , Xiangyu Peng , Yilun Zhou , Yixin Mao , Vaibhav Vats , Linnea Ross , Divyansh Agarwal , Caiming Xiong , Chien-Sheng Wu

Large Language Models (LLMs) are powerful tools for answering user queries, yet they remain highly vulnerable to jailbreak attacks. Existing guardrail methods typically rely on internal features or textual responses to detect malicious…

Cryptography and Security · Computer Science 2026-05-29 Zikai Zhang , Rui Hu , Olivera Kotevska , Jiahao Xu

Large Language Models (LLMs) face a significant threat from multi-turn jailbreak attacks, where adversaries progressively steer conversations to elicit harmful outputs. However, the practical effectiveness of existing attacks is undermined…

Cryptography and Security · Computer Science 2026-01-12 Songze Li , Ruishi He , Xiaojun Jia , Jun Wang , Zhihui Fu

Hidden malicious intent in multi-turn dialogue poses a growing threat to deployed large language models (LLMs). Rather than exposing a harmful objective in a single prompt, increasingly capable attackers can distribute their intent across…

Computation and Language · Computer Science 2026-05-13 Xinjie Shen , Rongzhe Wei , Peizhi Niu , Haoyu Wang , Ruihan Wu , Eli Chien , Bo Li , Pin-Yu Chen , Pan Li

Large language models (LLMs) are increasingly used for mental-health support; yet prevailing evaluation methods--fluency metrics, preference tests, and generic dialogue benchmarks--fail to capture the clinically critical dimensions of…

Computation and Language · Computer Science 2026-03-20 Fangrui Huang , Souhad Chbeir , Arpandeep Khatua , Sheng Wang , Sijun Tan , Kenan Ye , Lily Bailey , Merryn Daniel , Ryan Louie , Sanmi Koyejo , Ehsan Adeli

As Large Language Models (LLMs) and generative AI become increasingly widespread, concerns about content safety have grown in parallel. Currently, there is a clear lack of high-quality, human-annotated datasets that address the full…

While large language model-based agents demonstrate great potential in collaborative tasks, their interactivity also introduces security vulnerabilities. In this paper, we propose and model group collusive attacks, a highly destructive…

Artificial Intelligence · Computer Science 2026-03-17 Yiling Tao , Xinran Zheng , Shuo Yang , Meiling Tao , Xingjun Wang

Large Language Models (LLMs) have been demonstrated to generate illegal or unethical responses, particularly when subjected to "jailbreak." Research on jailbreak has highlighted the safety issues of LLMs. However, prior studies have…

Computation and Language · Computer Science 2024-10-31 Zhenhong Zhou , Jiuyang Xiang , Haopeng Chen , Quan Liu , Zherui Li , Sen Su

As older adults increasingly use LLM-based chatbots for companionship and assistance, a safety gap is emerging. Older adults may face vulnerabilities from social isolation, limited digital literacy, and cognitive decline, yet existing…

Human-Computer Interaction · Computer Science 2026-05-21 Changxuan Fan , Xi Yang , Yueyuan Zheng , Bin Zhou , Yuanping Wang , Wenbin Hu , Huihao Jing , Ki Sen Hung , Dazhao Du , Haoran Li , Janet Hui-wen Hsiao , Yangqiu Song

With the widespread application of Large Language Models (LLMs), their associated security issues have become increasingly prominent, severely constraining their trustworthy deployment in critical domains. This paper proposes a novel safety…

Artificial Intelligence · Computer Science 2025-11-18 Qi Li , Jianjun Xu , Pingtao Wei , Jiu Li , Peiqiang Zhao , Jiwei Shi , Xuan Zhang , Yanhui Yang , Xiaodong Hui , Peng Xu , Wenqin Shao

The deployment of Large Reasoning Models (LRMs) in high-stakes decision-making pipelines has introduced a novel and opaque attack surface: reasoning backdoors. In these attacks, the model's intermediate Chain-of-Thought (CoT) is manipulated…

Cryptography and Security · Computer Science 2026-03-04 Zhen Guo , Shanghao Shi , Hao Li , Shamim Yazdani , Ning Zhang , Reza Tourani

Mental health disorders are among the most prevalent diseases worldwide, affecting nearly one in four people. Despite their widespread impact, the intervention rate remains below 25%, largely due to the significant cooperation required from…

Computation and Language · Computer Science 2024-09-17 Sijie Ji , Xinzhe Zheng , Jiawei Sun , Renqi Chen , Wei Gao , Mani Srivastava

The rapid development of Multimodal Large Reasoning Models (MLRMs) has demonstrated broad application potential, yet their safety and reliability remain critical concerns that require systematic exploration. To address this gap, we conduct…

Computation and Language · Computer Science 2025-10-14 Xinyue Lou , You Li , Jinan Xu , Xiangyu Shi , Chi Chen , Kaiyu Huang

Large language models have seen widespread adoption, yet they remain vulnerable to multi-turn jailbreak attacks, threatening their safe deployment. This has led to the task of training automated multi-turn attackers to probe model safety…

Artificial Intelligence · Computer Science 2026-04-22 Xiqiao Xiong , Ouxiang Li , Zhuo Liu , Moxin Li , Wentao Shi , Fengbin Zhu , Qifan Wang , Fuli Feng

Recent advances in large language models (LLMs) have led to the development of powerful AI chatbots capable of engaging in natural and human-like conversations. However, these chatbots can be potentially harmful, exhibiting manipulative,…

Artificial Intelligence · Computer Science 2023-04-04 Baihan Lin , Djallel Bouneffouf , Guillermo Cecchi , Kush R. Varshney

We present MultiBreak, a scalable and diverse multi-turn jailbreak benchmark to evaluate large language model (LLM) safety. Multi-turn jailbreaks mimic natural conversational settings, making them easier to bypass safety-aligned LLM than…

Computation and Language · Computer Science 2026-05-05 Jialin Song , Xiaodong Liu , Weiwei Yang , Wuyang Chen , Mingqian Feng , Xuekai Zhu , Jianfeng Gao

We introduce findings and methods to facilitate evidence-based discussion about how large language models (LLMs) should behave in response to user signals of risk of suicidal thoughts and behaviors (STB). People are already using LLMs as…

Human-Computer Interaction · Computer Science 2025-11-03 Nick Judd , Alexandre Vaz , Kevin Paeth , Layla Inés Davis , Milena Esherick , Jason Brand , Inês Amaro , Tony Rousmaniere

Language models exhibit human-like cognitive vulnerabilities, such as emotional framing, that escape traditional behavioral alignment. We present CCS-7 (Cognitive Cybersecurity Suite), a taxonomy of seven vulnerabilities grounded in human…

Cryptography and Security · Computer Science 2025-08-15 Yuksel Aydin

As large language models (LLMs) increasingly mediate emotionally sensitive conversations, especially in mental health contexts, their ability to recognize and respond to high-risk situations becomes a matter of public safety. This study…

‹ Prev 1 3 4 5 6 7 10 Next ›