English
Related papers

Related papers: $R^2$-Guard: Robust Reasoning Enabled LLM Guardrai…

200 papers

Large Language Models (LLMs) are susceptible to adversarial attacks such as jailbreaking, which can elicit harmful or unsafe behaviors. This vulnerability is exacerbated in multilingual settings, where multilingual safety-aligned data is…

Computation and Language · Computer Science 2025-09-29 Yahan Yang , Soham Dan , Shuo Li , Dan Roth , Insup Lee

Although the integration of large language models (LLMs) into robotics has unlocked transformative capabilities, it has also introduced significant safety concerns, ranging from average-case LLM errors (e.g., hallucinations) to adversarial…

Robotics · Computer Science 2026-03-05 Zachary Ravichandran , Alexander Robey , Vijay Kumar , George J. Pappas , Hamed Hassani

Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrails typically rely on single-pass classification or, more recently, distilled reasoning.…

Artificial Intelligence · Computer Science 2026-05-29 Siddharth Sai , Xiaofei Wen , Muhao Chen

Reasoning-based language models have demonstrated strong performance across various domains, with the most notable gains seen in mathematical and coding tasks. Recent research has shown that reasoning also offers significant benefits for…

Artificial Intelligence · Computer Science 2025-05-27 Makesh Narsimhan Sreedhar , Traian Rebedea , Christopher Parisien

As LLMs increasingly impact safety-critical applications, ensuring their safety using guardrails remains a key challenge. This paper proposes GuardReasoner, a new safeguard for LLMs, by guiding the guard model to learn to reason.…

Cryptography and Security · Computer Science 2025-10-20 Yue Liu , Hongcheng Gao , Shengfang Zhai , Yufei He , Jun Xia , Zhengyu Hu , Yulin Chen , Xihong Yang , Jiaheng Zhang , Stan Z. Li , Hui Xiong , Bryan Hooi

Large Language Models (LLMs) continue to exhibit vulnerabilities despite deliberate safety alignment efforts, posing significant risks to users and society. To safeguard against the risk of policy-violating content, system-level moderation…

Artificial Intelligence · Computer Science 2025-10-27 Jingnan Zheng , Xiangtian Ji , Yijun Lu , Chenhang Cui , Weixiang Zhao , Gelei Deng , Zhenkai Liang , An Zhang , Tat-Seng Chua

Ensuring the safety of large language models (LLMs) is critical as they are deployed in real-world applications. Existing guardrails rely on rule-based filtering or single-pass classification, limiting their ability to handle nuanced safety…

Computation and Language · Computer Science 2025-05-29 Xiaofei Wen , Wenxuan Zhou , Wenjie Jacky Mo , Muhao Chen

To enhance the safety of VLMs, this paper introduces a novel reasoning-based VLM guard model dubbed GuardReasoner-VL. The core idea is to incentivize the guard model to deliberatively reason before making moderation decisions via online RL.…

Artificial Intelligence · Computer Science 2025-05-19 Yue Liu , Shengfang Zhai , Mingzhe Du , Yulin Chen , Tri Cao , Hongcheng Gao , Cheng Wang , Xinfeng Li , Kun Wang , Junfeng Fang , Jiaheng Zhang , Bryan Hooi

The emergence of Large Reasoning Models (LRMs) introduces a new paradigm of explicit reasoning, enabling remarkable advances yet posing unique risks such as reasoning manipulation and information leakage. To mitigate these risks, current…

Artificial Intelligence · Computer Science 2026-02-03 Jingnan Zheng , Jingjun Xu , Yanzhen Luo , Chenhang Cui , Gelei Deng , Zhenkai Liang , Xiang Wang , An Zhang , Tat-Seng Chua

Large Reasoning Models (LRMs) have recently demonstrated impressive performances across diverse domains. However, how the safety of Large Language Models (LLMs) benefits from enhanced reasoning capabilities against jailbreak queries remains…

Computation and Language · Computer Science 2025-09-23 Junda Zhu , Lingyong Yan , Shuaiqiang Wang , Dawei Yin , Lei Sha

Large Reasoning Models (LRMs) have demonstrated impressive performance in reasoning-intensive tasks, but they remain vulnerable to harmful content generation, particularly in the mid-to-late steps of their reasoning processes. Current…

Computation and Language · Computer Science 2026-05-07 Yuquan Wang , Mi Zhang , Yining Wang , Geng Hong , Mi Wen , Xiaoyu You , Min Yang

Recent reasoning-based safety guardrails for Large Reasoning Models (LRMs), such as deliberative alignment, have shown strong defense against jailbreak attacks. By leveraging LRMs' reasoning ability, these guardrails help the models to…

Cryptography and Security · Computer Science 2025-10-24 Shuo Chen , Zhen Han , Haokun Chen , Bailan He , Shengyun Si , Jingpei Wu , Philip Torr , Volker Tresp , Jindong Gu

As LLMs become widespread across diverse applications, concerns about the security and safety of LLM interactions have intensified. Numerous guardrail models and benchmarks have been developed to ensure LLM content safety. However, existing…

Cryptography and Security · Computer Science 2026-02-13 Mintong Kang , Zhaorun Chen , Chejian Xu , Jiawei Zhang , Chengquan Guo , Minzhou Pan , Ivan Revilla , Yu Sun , Bo Li

The rapid advancement of large language models (LLMs) has increased the need for guardrail models to ensure responsible use, particularly in detecting unsafe and illegal content. While substantial safety data exist in English, multilingual…

Computation and Language · Computer Science 2025-02-10 Yihe Deng , Yu Yang , Junkai Zhang , Wei Wang , Bo Li

Aligning Vision-Language Models (VLMs) with safety standards is essential to mitigate risks arising from their multimodal complexity, where integrating vision and language unveils subtle threats beyond the reach of conventional safeguards.…

Machine Learning · Computer Science 2025-10-14 Menglan Chen , Xianghe Pang , Jingjing Dong , WenHao Wang , Yaxin Du , Siheng Chen

With the increasing adoption of large language models (LLMs), ensuring the safety of LLM systems has become a pressing concern. External LLM-based guardrail models have emerged as a popular solution to screen unsafe inputs and outputs, but…

Computation and Language · Computer Science 2025-10-08 Yining She , Daniel W. Peterson , Marianne Menglin Liu , Vikas Upadhyay , Mohammad Hossein Chaghazardi , Eunsuk Kang , Dan Roth

While Multimodal Large Language Models (MLLMs) have made remarkable progress in vision-language reasoning, they are also more susceptible to producing harmful content compared to models that focus solely on text. Existing defensive…

Computation and Language · Computer Science 2024-12-30 Yilei Jiang , Yingshui Tan , Xiangyu Yue

Guardrails are critical for the safe deployment of Large Language Models (LLMs)-powered software. Unlike traditional rule-based systems with limited, predefined input-output spaces that inherently constrain unsafe behavior, LLMs enable…

Cryptography and Security · Computer Science 2025-09-23 Rui Yang , Michael Fu , Chakkrit Tantithamthavorn , Chetan Arora , Gunel Gulmammadova , Joey Chua

Guardrails have emerged as an alternative to safety alignment for content moderation of large language models (LLMs). Existing model-based guardrails have not been designed for resource-constrained computational portable devices, such as…

Machine Learning · Computer Science 2024-12-19 Hayder Elesedy , Pedro M. Esperança , Silviu Vlad Oprea , Mete Ozay

The rapid development of Multimodal Large Reasoning Models (MLRMs) has demonstrated broad application potential, yet their safety and reliability remain critical concerns that require systematic exploration. To address this gap, we conduct…

Computation and Language · Computer Science 2025-10-14 Xinyue Lou , You Li , Jinan Xu , Xiangyu Shi , Chi Chen , Kaiyu Huang
‹ Prev 1 2 3 10 Next ›