English
Related papers

Related papers: Ablating Safety: Mechanisms for Removing Alignment…

200 papers

Access control is a cornerstone of secure computing, yet large language models often blur role boundaries by producing unrestricted responses. We study role-conditioned refusals, focusing on the LLM's ability to adhere to access control…

Computation and Language · Computer Science 2025-10-10 Đorđe Klisura , Joseph Khoury , Ashish Kundu , Ram Krishnan , Anthony Rios

Safety fine-tuning of language models typically requires a curated adversarial dataset. We take a different approach: score each candidate prompt's difficulty by how often the target model's own rollouts are judged harmful, then fine-tune…

Machine Learning · Computer Science 2026-05-06 Prakhar Gupta , Garv Shah , Donghua Zhang

The safety alignment ability of Vision-Language Models (VLMs) is prone to be degraded by the integration of the vision module compared to its LLM backbone. We investigate this phenomenon, dubbed as ''safety alignment degradation'' in this…

Computation and Language · Computer Science 2024-10-14 Qin Liu , Chao Shang , Ling Liu , Nikolaos Pappas , Jie Ma , Neha Anna John , Srikanth Doss , Lluis Marquez , Miguel Ballesteros , Yassine Benajiba

In perpetrator treatment, a recurring observation is the dissociation between insight and action: offenders articulate remorse yet behavioral change does not follow. We report four preregistered studies (1,584 multi-agent simulations across…

Artificial Intelligence · Computer Science 2026-03-06 Hiroki Fukui

Safety-aligned large language models rely on RLHF and instruction tuning to refuse harmful requests, yet the internal mechanisms implementing safety behavior remain poorly understood. We introduce the Attention Redistribution Attack (ARA),…

Cryptography and Security · Computer Science 2026-05-04 Aviral Srivastava , Sourav Panda

Recent studies reveal that integrating new modalities into Large Language Models (LLMs), such as Vision-Language Models (VLMs), creates a new attack surface that bypasses existing safety training techniques like Supervised Fine-tuning (SFT)…

Computation and Language · Computer Science 2025-10-15 Trishna Chakraborty , Erfan Shayegani , Zikui Cai , Nael Abu-Ghazaleh , M. Salman Asif , Yue Dong , Amit K. Roy-Chowdhury , Chengyu Song

Hallucination in large language models (LLMs) has been widely studied in recent years, with progress in both detection and mitigation aimed at improving truthfulness. Yet, a critical side effect remains largely overlooked: enhancing…

Computation and Language · Computer Science 2026-02-02 Omar Mahmoud , Ali Khalil , Buddhika Laknath Semage , Thommen George Karimpanal , Santu Rana

Safety-aligned language models often refuse prompts that are actually harmless. Current evaluations mostly report global rates such as false rejection or compliance. These scores treat each prompt alone and miss local inconsistency, where a…

Computation and Language · Computer Science 2025-12-22 Riad Ahmed Anonto , Md Labid Al Nahiyan , Md Tanvir Hassan

Current safety evaluations of language models rely on benchmark-based assessments that may miss localized vulnerabilities. We present RepIt, a simple and data-efficient framework for isolating concept-specific representations in LM…

Artificial Intelligence · Computer Science 2026-04-22 Vincent Siu , Nathan W. Henry , Nicholas Crispino , Yang Liu , Dawn Song , Chenguang Wang

Large Language and Vision-Language Models (LLMs/VLMs) are increasingly used in safety-critical applications, yet their opaque decision-making complicates risk assessment and reliability. Uncertainty quantification (UQ) helps assess…

Machine Learning · Computer Science 2025-02-12 Sina Tayebati , Divake Kumar , Nastaran Darabi , Dinithi Jayasuriya , Ranganath Krishnan , Amit Ranjan Trivedi

Large Language Models (LLMs) often exhibit significant behavioral shifts when they perceive a change from a real-world deployment context to a controlled evaluation setting, a phenomenon known as "evaluation awareness." This discrepancy…

Computation and Language · Computer Science 2025-12-05 Lang Xiong , Nishant Bhargava , Jianhang Hong , Jeremy Chang , Haihao Liu , Vasu Sharma , Kevin Zhu

Safety alignment of Large Language Models (LLMs) has recently become a critical objective of model developers. In response, a growing body of work has been investigating how safety alignment can be bypassed through various jailbreaking…

Machine Learning · Computer Science 2024-12-06 Jason Vega , Junsheng Huang , Gaokai Zhang , Hangoo Kang , Minjia Zhang , Gagandeep Singh

Safety benchmark scores provide incomplete evidence of deployment readiness: aligned language models often adhere to rigid rules even when a situational update flips which action is safe. We term this failure brittle safety. To diagnose it,…

Artificial Intelligence · Computer Science 2026-05-28 Dasol Choi , Alex Kwon

Safety evaluations of language models often treat serving configuration as fixed background infrastructure, but batch condition is an untested treatment variable whenever the same prompt may be evaluated alone, in a synchronized batch, or…

Machine Learning · Computer Science 2026-05-28 Sahil Kadadekar

Text-to-image diffusion models can emit copyrighted, unsafe, or private content. Safety alignment aims to suppress specific concepts, yet evaluations seldom test whether safety persists under benign downstream fine-tuning routinely applied…

Cryptography and Security · Computer Science 2025-11-26 Mohammed Talha Alam , Nada Saadi , Fahad Shamshad , Nils Lukas , Karthik Nandakumar , Fahkri Karray , Samuele Poppi

Large Language Models (LLMs) are widely used across sectors, yet their alignment with International Humanitarian Law (IHL) is not well understood. This study evaluates eight leading LLMs on their ability to refuse prompts that explicitly…

Computers and Society · Computer Science 2025-06-10 John Mavi , Diana Teodora Găitan , Sergio Coronado

Refusal mechanisms in large language models (LLMs) are essential for ensuring safety. Recent research has revealed that refusal behavior can be mediated by a single direction in activation space, enabling targeted interventions to bypass…

Computation and Language · Computer Science 2026-02-26 Xinpeng Wang , Mingyang Wang , Yihong Liu , Hinrich Schütze , Barbara Plank

Vision-language-action models (VLAs) have been extensively used in robotics applications, achieving great success in various manipulation problems. More recently, VLAs have been used in long-horizon tasks and evaluated on benchmarks, such…

Robotics · Computer Science 2026-04-24 Amir Rasouli , Yangzheng Wu , Zhiyuan Li , Rui Heng Yang , Xuan Zhao , Charles Eret , Sajjad Pakdamansavoji

Lifelong multimodal agents must continuously adapt to new tasks through post-training, but this creates a fundamental tension between acquiring capabilities and preserving safety alignment. We demonstrate that fine-tuning aligned…

Artificial Intelligence · Computer Science 2026-03-17 Idhant Gulati , Shivam Raval

Low-Rank Adaptation (LoRA) is widely used for parameter-efficient fine-tuning of large language models, but it is notably ineffective at removing backdoor behaviors from poisoned pretrained models when fine-tuning on clean dataset. Contrary…

Computation and Language · Computer Science 2026-01-13 Hoang-Chau Luong , Lingwei Chen
‹ Prev 1 3 4 5 6 7 10 Next ›