English
Related papers

Related papers: A Single Neuron Is Sufficient to Bypass Safety Ali…

200 papers

Large Language Models used in ChatGPT have traditionally been trained to learn a refusal boundary: depending on the user's intent, the model is taught to either fully comply or outright refuse. While this is a strong mitigation for…

Computers and Society · Computer Science 2025-08-14 Yuan Yuan , Tina Sriskandarajah , Anna-Luisa Brakman , Alec Helyar , Alex Beutel , Andrea Vallone , Saachi Jain

Large Language Models (LLMs) are aligned to meet ethical standards and safety requirements by training them to refuse answering harmful or unsafe prompts. In this paper, we demonstrate how adversaries can exploit LLMs' alignment to implant…

Machine Learning · Computer Science 2025-08-29 Md Abdullah Al Mamun , Ihsen Alouani , Nael Abu-Ghazaleh

Word alignments identify translational correspondences between words in a parallel sentence pair and is used, for instance, to learn bilingual dictionaries, to train statistical machine translation systems , or to perform quality…

Computation and Language · Computer Science 2020-09-29 Anh Khoa Ngo Ho , François Yvon

We show that in a drive-response coupling framework extreme events are suppressed in the response system by the dominance of a single driving signal. We validate this approach across three distinct response network topologies, namely (i) a…

Chaotic Dynamics · Physics 2026-02-03 R. Shashangan , S. Sudharsan , Dibakar Ghosh , M. Senthilvelan

We have witnessed superhuman intelligence thanks to the fast development of large language models and multimodal language models. As the application of such superhuman models becomes more and more popular, a critical question arises here:…

Computation and Language · Computer Science 2024-12-24 Minlie Huang , Yingkang Wang , Shiyao Cui , Pei Ke , Jie Tang

When LLMs are deployed in sensitive, human-facing settings, it is crucial that they do not output unsafe, biased, or privacy-violating outputs. For this reason, models are both trained and instructed to refuse to answer unsafe prompts such…

Machine Learning · Computer Science 2024-07-04 Leon Lin , Hannah Brown , Kenji Kawaguchi , Michael Shieh

Malfunctioning neurons in the brain sometimes operate synchronously, reportedly causing many neurological diseases, e.g. Parkinson's. Suppression and control of this collective synchronous activity are therefore of great importance for…

Neurons and Cognition · Quantitative Biology 2021-09-22 Dmitrii Krylov , Remi Tachet , Romain Laroche , Michael Rosenblum , Dmitry V. Dylov

Ensuring Large Language Model (LLM) safety remains challenging due to the absence of universal standards and reliable content validators, making it difficult to obtain effective training signals. We discover that aligned models already…

Artificial Intelligence · Computer Science 2025-10-02 Guobin Shen , Dongcheng Zhao , Haibo Tong , Jindong Li , Feifei Zhao , Yi Zeng

Emergence, the phenomenon of a rapid performance increase once the model scale reaches a threshold, has achieved widespread attention recently. The literature has observed that monosemantic neurons in neural networks gradually diminish as…

Emerging Technologies · Computer Science 2025-04-01 Jiachuan Wang , Shimin Di , Tianhao Tang , Haoyang LI , Charles Wang-wai Ng , Xiaofang Zhou , Lei Chen

There have been numerous advances in reinforcement learning, but the typically unconstrained exploration of the learning process prevents the adoption of these methods in many safety critical applications. Recent work in safe reinforcement…

Machine Learning · Computer Science 2019-10-02 David Isele , Alireza Nakhaei , Kikuo Fujimura

This paper presents a comprehensive empirical study on the safety alignment capabilities. We evaluate what matters for safety alignment in LLMs and LRMs to provide essential insights for developing more secure and reliable AI systems. We…

Computation and Language · Computer Science 2026-02-25 Xing Li , Hui-Ling Zhen , Lihao Yin , Xianzhi Yu , Zhenhua Dong , Mingxuan Yuan

Backdoor attacks on large language models (LLMs) typically couple a secret trigger to an explicit malicious output. We show that this explicit association is unnecessary for common LLMs. We introduce a compliance-only backdoor: supervised…

Machine Learning · Computer Science 2025-11-18 Yuting Tan , Yi Huang , Zhuo Li

Current safety alignment for large language models(LLMs) continues to present vulnerabilities, given that adversarial prompting can effectively bypass their safety measures.Our investigation shows that these safety mechanisms predominantly…

Cryptography and Security · Computer Science 2025-08-28 Chao Huang , Zefeng Zhang , Juewei Yue , Quangang Li , Chuang Zhang , Tingwen Liu

While there has been progress towards aligning Large Language Models (LLMs) with human values and ensuring safe behaviour at inference time, safety guards can easily be removed when fine tuned on unsafe and harmful datasets. While this…

Vision-Language Models (VLMs) have achieved remarkable progress in multimodal reasoning tasks through enhanced chain-of-thought capabilities. However, this advancement also introduces novel safety risks, as these models become increasingly…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Yinan Xia , Yilei Jiang , Yingshui Tan , Xiaoyong Zhu , Xiangyu Yue , Bo Zheng

This work addresses the computational challenge of enforcing privacy for agentic Large Language Models (LLMs), where privacy is governed by the contextual integrity framework. Indeed, existing defenses rely on LLM-mediated checking stages…

Cryptography and Security · Computer Science 2026-01-22 Saswat Das , Ferdinando Fioretto

Safety-oriented instruction-following is supposed to keep LLM-controlled robots safe. We show it also creates an availability attack surface. By injecting short safety-plausible phrases (1-5 tokens) into a robots audio channel, an adversary…

Cryptography and Security · Computer Science 2026-04-29 Jonathan Steinberg , Oren Gal

Safety benchmark scores provide incomplete evidence of deployment readiness: aligned language models often adhere to rigid rules even when a situational update flips which action is safe. We term this failure brittle safety. To diagnose it,…

Artificial Intelligence · Computer Science 2026-05-28 Dasol Choi , Alex Kwon

Large language model (LLM)-based conversational AI systems present a challenge to human cognition that current frameworks for understanding misinformation and persuasion do not adequately address. This paper proposes that a significant…

Human-Computer Interaction · Computer Science 2026-05-27 Andrew D. Maynard

As large-scale language models increasingly impact safety-critical domains, ensuring their reliable adherence to well-defined principles remains a fundamental challenge. We introduce Deliberative Alignment, a new paradigm that directly…

‹ Prev 1 8 9 10 Next ›