English
Related papers

Related papers: SafeNeuron: Neuron-Level Safety Alignment for Larg…

200 papers

The rapid advancement of large language models (LLMs) has demonstrated milestone success in a variety of tasks, yet their potential for generating harmful content has raised significant safety concerns. Existing safety evaluation approaches…

Computation and Language · Computer Science 2025-05-22 Tianqi Du , Zeming Wei , Quan Chen , Chenheng Zhang , Yisen Wang

Embodied agents powered by large language models (LLMs) inherit advanced planning capabilities; however, their direct interaction with the physical world exposes them to safety vulnerabilities. In this work, we identify four key reasoning…

Artificial Intelligence · Computer Science 2025-10-01 Ruolin Chen , Yinqian Sun , Jihang Wang , Mingyang Lv , Qian Zhang , Yi Zeng

Large Language Models (LLMs) are vulnerable to jailbreak attacks that exploit weaknesses in traditional safety alignment, which often relies on rigid refusal heuristics or representation engineering to block harmful outputs. While they are…

Computation and Language · Computer Science 2025-10-01 Yuyou Zhang , Miao Li , William Han , Yihang Yao , Zhepeng Cen , Ding Zhao

Large Reasoning Models (LRMs) have significantly improved problem-solving through explicit Chain-of-Thought (CoT) reasoning. However, this capability creates a Safety-Helpfulness Paradox: the reasoning process itself can be misused to…

Artificial Intelligence · Computer Science 2026-01-27 Xin Gao , Shaohan Yu , Zerui Chen , Yueming Lyu , Weichen Yu , Guanghao Li , Jiyao Liu , Jianxiong Gao , Jian Liang , Ziwei Liu , Chenyang Si

We introduce SafeWork-R1, a cutting-edge multimodal reasoning model that demonstrates the coevolution of capabilities and safety. It is developed by our proposed SafeLadder framework, which incorporates large-scale, progressive,…

Artificial Intelligence · Computer Science 2025-08-08 Shanghai AI Lab , : , Yicheng Bao , Guanxu Chen , Mingkang Chen , Yunhao Chen , Chiyu Chen , Lingjie Chen , Sirui Chen , Xinquan Chen , Jie Cheng , Yu Cheng , Dengke Deng , Yizhuo Ding , Dan Ding , Xiaoshan Ding , Yi Ding , Zhichen Dong , Lingxiao Du , Yuyu Fan , Xinshun Feng , Yanwei Fu , Yuxuan Gao , Ruijun Ge , Tianle Gu , Lujun Gui , Jiaxuan Guo , Qianxi He , Yuenan Hou , Xuhao Hu , Hong Huang , Kaichen Huang , Shiyang Huang , Yuxian Jiang , Shanzhe Lei , Jie Li , Lijun Li , Hao Li , Juncheng Li , Xiangtian Li , Yafu Li , Lingyu Li , Xueyan Li , Haotian Liang , Dongrui Liu , Qihua Liu , Zhixuan Liu , Bangwei Liu , Huacan Liu , Yuexiao Liu , Zongkai Liu , Chaochao Lu , Yudong Lu , Xiaoya Lu , Zhenghao Lu , Qitan Lv , Caoyuan Ma , Jiachen Ma , Xiaoya Ma , Zhongtian Ma , Lingyu Meng , Ziqi Miao , Yazhe Niu , Yuezhang Peng , Yuan Pu , Han Qi , Chen Qian , Xingge Qiao , Jingjing Qu , Jiashu Qu , Wanying Qu , Wenwen Qu , Xiaoye Qu , Qihan Ren , Qingnan Ren , Qingyu Ren , Jing Shao , Wenqi Shao , Shuai Shao , Dongxing Shi , Xin Song , Xinhao Song , Yan Teng , Xuan Tong , Yingchun Wang , Xuhong Wang , Shujie Wang , Xin Wang , Yige Wang , Yixu Wang , Yuanfu Wang , Futing Wang , Ruofan Wang , Wenjie Wang , Yajie Wang , Muhao Wei , Xiaoyu Wen , Fenghua Weng , Yuqi Wu , Yingtong Xiong , Xingcheng Xu , Chao Yang , Yue Yang , Yang Yao , Yulei Ye , Zhenyun Yin , Yi Yu , Bo Zhang , Qiaosheng Zhang , Jinxuan Zhang , Yexin Zhang , Yinqiang Zheng , Hefeng Zhou , Zhanhui Zhou , Pengyu Zhu , Qingzi Zhu , Yubo Zhu , Bowen Zhou

The growing awareness of safety concerns in large language models (LLMs) has sparked considerable interest in the evaluation of safety. This study investigates an under-explored issue about the evaluation of LLMs, namely the substantial…

Computation and Language · Computer Science 2024-04-02 Yixu Wang , Yan Teng , Kexin Huang , Chengqi Lyu , Songyang Zhang , Wenwei Zhang , Xingjun Ma , Yu-Gang Jiang , Yu Qiao , Yingchun Wang

Federated learning (FL) addresses privacy and data-silo issues in the training of large language models (LLMs). Most prior work focuses on improving the efficiency of federated learning for LLMs (FedLLM). However, security in open federated…

Cryptography and Security · Computer Science 2026-04-21 Mingxiang Tao , Yu Tian , Wenxuan Tu , Yue Yang , Xue Yang , Xiangyan Tang

Large language models (LLMs) are vulnerable when trained on datasets containing harmful content, which leads to potential jailbreaking attacks in two scenarios: the integration of harmful texts within crowdsourced data used for pre-training…

Cryptography and Security · Computer Science 2024-06-03 Xiaoqun Liu , Jiacheng Liang , Muchao Ye , Zhaohan Xi

Large Language Models (LLMs) have advanced various Natural Language Processing (NLP) tasks, such as text generation and translation, among others. However, these models often generate texts that can perpetuate biases. Existing approaches to…

Computation and Language · Computer Science 2025-01-07 Shaina Raza , Oluwanifemi Bamgbose , Shardul Ghuge , Fatemeh Tavakol , Deepak John Reji , Syed Raza Bashir

Large language models (LLMs) rely on safety alignment to avoid responding to malicious user inputs. Unfortunately, jailbreak can circumvent safety guardrails, resulting in LLMs generating harmful content and raising concerns about LLM…

Computation and Language · Computer Science 2024-06-14 Zhenhong Zhou , Haiyang Yu , Xinghua Zhang , Rongwu Xu , Fei Huang , Yongbin Li

When building Large Language Models (LLMs), it is paramount to bear safety in mind and protect them with guardrails. Indeed, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may…

Computation and Language · Computer Science 2024-06-25 Simone Tedeschi , Felix Friedrich , Patrick Schramowski , Kristian Kersting , Roberto Navigli , Huu Nguyen , Bo Li

Deploying large language models (LLMs) in real-world applications requires robust safety guard models to detect and block harmful user prompts. While large safety guard models achieve strong performance, their computational cost is…

Computation and Language · Computer Science 2025-05-23 Seanie Lee , Dong Bok Lee , Dominik Wagner , Minki Kang , Haebin Seong , Tobias Bocklet , Juho Lee , Sung Ju Hwang

With the rapid development of large language models (LLMs), they are not only used as general-purpose AI assistants but are also customized through further fine-tuning to meet the requirements of different applications. A pivotal factor in…

Computation and Language · Computer Science 2024-01-23 Pengyu Wang , Dong Zhang , Linyang Li , Chenkun Tan , Xinghao Wang , Ke Ren , Botian Jiang , Xipeng Qiu

The current paradigm for safety alignment of large language models (LLMs) follows a one-size-fits-all approach: the model refuses to interact with any content deemed unsafe by the model provider. This approach lacks flexibility in the face…

Computation and Language · Computer Science 2025-03-05 Jingyu Zhang , Ahmed Elgohary , Ahmed Magooda , Daniel Khashabi , Benjamin Van Durme

Multi-turn jailbreak attacks progressively erode LLM safety alignment across seemingly innocuous conversation turns, achieving success rates exceeding 90% against state-of-the-art models. Existing alignment-based and guardrail methods…

Cryptography and Security · Computer Science 2026-04-21 Bo Yan , Weikai Lin , Yada Zhu , Song Wang

This work addresses the computational challenge of enforcing privacy for agentic Large Language Models (LLMs), where privacy is governed by the contextual integrity framework. Indeed, existing defenses rely on LLM-mediated checking stages…

Cryptography and Security · Computer Science 2026-01-22 Saswat Das , Ferdinando Fioretto

In the rapidly evolving field of Large Language Models (LLMs), ensuring safety is a crucial and widely discussed topic. However, existing works often overlook the geo-diversity of cultural and legal standards across the world. To…

Computation and Language · Computer Science 2024-12-10 Da Yin , Haoyi Qiu , Kung-Hsiang Huang , Kai-Wei Chang , Nanyun Peng

Over the last decade, Neural Networks (NNs) have been widely used in numerous applications including safety-critical ones such as autonomous systems. Despite their emerging adoption, it is well known that NNs are susceptible to Adversarial…

Machine Learning · Computer Science 2022-07-19 Dor Cohen , Ofer Strichman

Large language models (LLMs) have brought significant advancements to code generation, benefiting both novice and experienced developers. However, their training using unsanitized data from open-source repositories, like GitHub, introduces…

Software Engineering · Computer Science 2023-10-26 Jiexin Wang , Liuwen Cao , Xitong Luo , Zhiping Zhou , Jiayuan Xie , Adam Jatowt , Yi Cai

The progress of AI systems such as large language models (LLMs) raises increasingly pressing concerns about their safe deployment. This paper examines the value alignment problem for LLMs, arguing that current alignment strategies are…

Computation and Language · Computer Science 2025-06-06 Raphaël Millière