English
Related papers

Related papers: Refusal Before Decoding: Detecting and Exploiting …

200 papers

LLMs utilizing chain-of-thought reasoning often waste substantial compute by producing long, incorrect responses. Abstention can mitigate this by withholding outputs unlikely to be correct. While most abstention methods decide to withhold…

Machine Learning · Computer Science 2026-05-26 Hen Davidov , Nachshon Cohen , Oren Kalinsky , Yaron Fairstein , Guy Kushilevitz , Ram Yazdi , Patrick Rebeschini

Prior behavioural work suggests that some LLMs alter choices when options are framed as causing pain or pleasure, and that such deviations can scale with stated intensity. To bridge behavioural evidence (what the model does) with…

Artificial Intelligence · Computer Science 2026-02-24 Francesca Bianco , Derek Shiller

Latent Action Models (LAMs) have rapidly gained traction as an important component in the pre-training pipelines of leading Vision-Language-Action models. However, they fail when observations contain action-correlated distractors, often…

Instruction Fine-Tuning (IFT) has been widely adopted as an effective post-training strategy to enhance various abilities of Large Language Models (LLMs). However, prior studies have shown that IFT can significantly compromise LLMs' safety,…

Computation and Language · Computer Science 2025-09-09 Yanrui Du , Fenglei Fan , Sendong Zhao , Jiawei Cao , Qika Lin , Kai He , Ting Liu , Bing Qin , Mengling Feng

Reinforcement Learning (RL) significantly enhances the reasoning abilities of large language models (LLMs), yet applying it to multi-turn agentic tasks remains challenging due to the long-horizon nature of interactions and the stochasticity…

Artificial Intelligence · Computer Science 2026-04-03 Jingyue Gao , Yanjiang Guo , Xiaoshuai Chen , Jianyu Chen

Jailbreak attacks pose persistent threats to large language models (LLMs). Current safety alignment methods have attempted to address these issues, but they experience two significant limitations: insufficient safety alignment depth and…

Cryptography and Security · Computer Science 2025-09-19 Yuanbo Xie , Yingjie Zhang , Tianyun Liu , Duohe Ma , Tingwen Liu

Current answering paradigms for Large Reasoning Models (LRMs) often fail to account for the fact that some questions may lie beyond the model's operational capability boundary, leading to long but unproductive reasoning. In this paper, we…

Artificial Intelligence · Computer Science 2026-03-17 Qingjie Zhang , Yujia Fu , Yang Wang , Liu Yan , Tao Wei , Ke Xu , Minlie Huang , Han Qiu

Large Language Model (LLM)-powered agents demonstrate strong capabilities in autonomous task execution, tool use, and multi-step reasoning. However, their increasing autonomy also introduces a new attack surface: adversarial interactions…

Artificial Intelligence · Computer Science 2026-05-05 Sheldon Yu , Yingcheng Sun , Hanqing Guo , Julian McAuley , Qianqian Tong

Refusals - instances where large language models (LLMs) decline or fail to fully execute user instructions - are crucial for both AI safety and AI capabilities and the reduction of hallucinations in particular. These behaviors are learned…

Artificial Intelligence · Computer Science 2024-12-24 Alexander von Recum , Christoph Schnabl , Gabor Hollbeck , Silas Alberti , Philip Blinde , Marvin von Hagen

We introduce a method to reduce refusal rates of large language models (LLMs) on sensitive content without modifying model weights or prompts. Motivated by the observation that refusals in certain models were often preceded by the specific…

Computation and Language · Computer Science 2025-06-02 Harvey Dam , Jonas Knochelmann , Vinu Joseph , Ganesh Gopalakrishnan

Language models are instruction-tuned to refuse harmful requests, but the mechanisms underlying this behavior remain poorly understood. Popular steering methods operate on the residual stream and degrade output coherence at high…

Machine Learning · Computer Science 2026-05-13 Sam Herring , Jake Naviasky , Karan Malhotra

Large Language Models (LLMs) encode behaviors such as refusal within their activation space, yet identifying these behaviors remains a significant challenge. Existing methods often rely on predefined refusal templates detectable in output…

Computation and Language · Computer Science 2025-06-03 Vincent Siu , Nicholas Crispino , Zihao Yu , Sam Pan , Zhun Wang , Yang Liu , Dawn Song , Chenguang Wang

Tool-using LLM agents produce trajectories whose calls form a directed dependency graph: earlier tool outputs supply arguments to later calls. Whether this execution structure is represented inside the model is unknown; prior structural…

Computation and Language · Computer Science 2026-05-26 Tianda Sun , Dimitar Kazakov

Most jailbreak techniques for Large Language Models (LLMs) primarily rely on prompt modifications, including paraphrasing, obfuscation, or conversational strategies. Meanwhile, abliteration techniques (also known as targeted ablations of…

Cryptography and Security · Computer Science 2026-03-17 Maël Jenny , Jérémie Dentan , Sonia Vanier , Michaël Krajecki

Large Vision-Language Models (LVLMs) have shown remarkable capabilities across a wide range of multimodal tasks. However, their integration of visual inputs introduces expanded attack surfaces, thereby exposing them to novel security…

Computation and Language · Computer Science 2025-05-29 Juan Ren , Mark Dras , Usman Naseem

As LLMs are increasingly deployed in real-world applications, ensuring their ability to refuse malicious prompts, especially jailbreak attacks, is essential for safe and reliable use. Recently, activation steering has emerged as an…

Machine Learning · Computer Science 2026-02-10 Leheng Sheng , Changshuo Shen , Weixiang Zhao , Junfeng Fang , Xiaohao Liu , Zhenkai Liang , Xiang Wang , An Zhang , Tat-Seng Chua

Prior work argues that refusal in large language models is mediated by a single activation-space direction, enabling effective steering and ablation. We show that this account is incomplete. Across eleven categories of refusal and…

Computation and Language · Computer Science 2026-02-03 Faaiz Joad , Majd Hawasly , Sabri Boughorbel , Nadir Durrani , Husrev Taha Sencar

Refusal mechanisms in large language models (LLMs) are essential for ensuring safety. Recent research has revealed that refusal behavior can be mediated by a single direction in activation space, enabling targeted interventions to bypass…

Computation and Language · Computer Science 2026-02-26 Xinpeng Wang , Mingyang Wang , Yihong Liu , Hinrich Schütze , Barbara Plank

Despite investments in improving model safety, studies show that misaligned capabilities remain latent in safety-tuned models. In this work, we shed light on the mechanics of this phenomenon. First, we show that even when model generations…

Computation and Language · Computer Science 2024-08-14 Asma Ghandeharioun , Ann Yuan , Marius Guerard , Emily Reif , Michael A. Lepori , Lucas Dixon

As LLM-based agents increasingly operate in multi-agent systems, understanding adversarial manipulation becomes critical for defensive design. We present a systematic study of intentional deception as an engineered capability, using…

Artificial Intelligence · Computer Science 2026-03-10 Jason Starace , Terence Soule