English
Related papers

Related papers: No Free Lunch with Guardrails

200 papers

There has been rapid development in generative AI tools across the education sector, which in turn is leading to increased adoption by teachers. However, this raises concerns regarding the safety and age-appropriateness of the AI-generated…

Computers and Society · Computer Science 2025-08-08 Hannah-Beth Clark , Laura Benton , Emma Searle , Margaux Dowland , Matthew Gregory , Will Gayne , John Roberts

Conversational AI systems require guardrails to prevent harmful outputs, yet existing approaches use static rules that cannot adapt to new threats or deployment contexts. We introduce Lattice, a framework for self-constructing and…

Artificial Intelligence · Computer Science 2026-01-27 Emily Broadhurst , Tawab Safi , Joseph Edell , Vashisht Ganesh , Karime Maamari

Safety guardrails have become an active area of research in AI safety, aimed at ensuring the appropriate behavior of large language models (LLMs). However, existing research lacks consideration of nuances across linguistic and cultural…

Cryptography and Security · Computer Science 2026-04-21 Hua-Rong Chu , Kuan-Chun Wang , Yao-Te Huang

Large Language Model (LLMs) such as ChatGPT that exhibit generative AI capabilities are facing accelerated adoption and innovation. The increased presence of Generative AI (GAI) inevitably raises concerns about the risks and safety…

Computers and Society · Computer Science 2024-07-29 Jaymari Chua , Yun Li , Shiyi Yang , Chen Wang , Lina Yao

Analyzing 500 CTF participants, this paper shows that while participants readily bypassed simple AI guardrails using common techniques, layered multi-step defenses still posed significant challenges, offering concrete insights for building…

Cryptography and Security · Computer Science 2025-10-21 Giacomo Bertollo , Naz Bodemir , Jonah Burgess

AI systems have found a wide range of real-world applications in recent years. The adoption of edge artificial intelligence, embedding AI directly into edge devices, is rapidly growing. Despite the implementation of guardrails and safety…

Hardware Architecture · Computer Science 2025-11-13 Eren Kurshan , Yuan Xie , Paul Franzon

Machine learning systems increasingly drive innovation across scientific fields and industry, yet challenges in compute overhead, specifically during inference, limit their scalability and sustainability. Responsible AI guardrails,…

Machine Learning · Computer Science 2024-12-30 Jason M. Pittman

We introduce a lightweight yet highly effective safety guardrail framework for language models, demonstrating that small-scale language models can achieve, and even surpass, the performance of larger counterparts in content moderation…

Machine Learning · Computer Science 2025-07-14 Aleksei Ilin , Gor Matevosyan , Xueying Ma , Vladimir Eremin , Suhaa Dada , Muqun Li , Riyaaz Shaik , Haluk Noyan Tokgozoglu

Jailbreak attacks bypass the guardrails of large language models to produce harmful outputs. In this paper, we ask whether the model outputs produced by existing jailbreaks are actually useful. For example, when jailbreaking a model to give…

Machine Learning · Computer Science 2025-04-16 Kristina Nikolić , Luze Sun , Jie Zhang , Florian Tramèr

Recent work has demonstrated that finetuning is a promising approach to 'unlearn' concepts from large language models. However, finetuning can be expensive, as it requires both generating a set of examples and running iterations of…

Computation and Language · Computer Science 2024-06-12 Pratiksha Thaker , Yash Maurya , Shengyuan Hu , Zhiwei Steven Wu , Virginia Smith

As AI agents move from chat interfaces to systems that read private data, call tools, and execute multi-step workflows, guardrails become a last line of defense against concrete deployment harms. In these settings, guardrail failures are no…

Large Language Models (LLMs) are typically aligned for safety during the post-training phase; however, they may still generate inappropriate outputs that could potentially pose risks to users. This challenge underscores the need for robust…

Machine Learning · Computer Science 2025-12-08 Mahesh Kumar Nandwana , Youngwan Lim , Joseph Liu , Alex Yang , Varun Notibala , Nishchaie Khanna

The security of Large Language Models (LLMs) has become an important research topic since the emergence of ChatGPT. Though there have been various effective methods to defend against jailbreak attacks, prefilling attacks remain an unsolved…

Cryptography and Security · Computer Science 2024-12-18 Zhiyu Xue , Guangliang Liu , Bocheng Chen , Kristen Marie Johnson , Ramtin Pedarsani

Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrails typically rely on single-pass classification or, more recently, distilled reasoning.…

Artificial Intelligence · Computer Science 2026-05-29 Siddharth Sai , Xiaofei Wen , Muhao Chen

Large Language Models (LLMs) deployed in production environments face a fundamental safety-utility trade-off either a strict filtering mechanisms prevent harmful outputs but often block benign queries or a relaxed controls risk unsafe…

Artificial Intelligence · Computer Science 2026-02-18 Ankit Sharma , Nachiket Tapas , Jyotiprakash Patra

Large Language Models (LLMs) are transforming enterprise workflows but introduce security and ethics challenges when employees inadvertently share confidential data or generate policy-violating content. This paper proposes SafeGPT, a…

Cryptography and Security · Computer Science 2026-05-26 Pratyush Desai , Luoxi Tang , Yuqiao Meng , Zhaohan Xi

In this paper, we show it is possible to bypass the safety guardrails of large language models (LLMs) through a humorous prompt including the unsafe request. In particular, our method does not edit the unsafe request and follows a fixed…

Computation and Language · Computer Science 2025-04-10 Pedro Cisneros-Velarde

The rise of Large Language Models has created a general excitement about the great potential for a myriad of applications. While LLMs offer many possibilities, questions about safety, privacy, and ethics have emerged, and all the key actors…

In this study, we disclose a worrying new vulnerability in Large Language Models (LLMs), which we term \textbf{involuntary jailbreak}. Unlike existing jailbreak attacks, this weakness is distinct in that it does not involve a specific…

Cryptography and Security · Computer Science 2025-12-30 Yangyang Guo , Yangyan Li , Mohan Kankanhalli
‹ Prev 1 3 4 5 6 7 10 Next ›