English
Related papers

Related papers: SaFeR-Steer: Evolving Multi-Turn MLLMs via Synthet…

200 papers

This paper introduces FRACTURED-SORRY-Bench, a framework for evaluating the safety of Large Language Models (LLMs) against multi-turn conversational attacks. Building upon the SORRY-Bench dataset, we propose a simple yet effective method…

Computation and Language · Computer Science 2024-11-08 Aman Priyanshu , Supriti Vijay

Large Multimodal Models (LMMs) have recently demonstrated their proficiency in holistic visual comprehension. However, most of them struggle to tackle region-level perception guided by visual prompts, especially for cases where multiple…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Yun Xing , Hanyuan Liu , Jiahao Nie , Shijian Lu

Activation steering has emerged as a powerful tool to shape LLM behavior without the need for weight updates. While its inherent brittleness and unreliability are well-documented, its safety implications remain underexplored. In this work,…

Cryptography and Security · Computer Science 2026-03-26 Yuxiao Li , Alina Fastowski , Efstratios Zaradoukas , Bardh Prenkaj , Gjergji Kasneci

Multimodal large language models (MLLMs) are essential for building general-purpose AI assistants; however, they pose increasing safety risks. How can we ensure safety alignment of MLLMs to prevent undesired behaviors? Going further, it is…

High-quality preference datasets are essential for training reward models that can effectively guide large language models (LLMs) in generating high-quality responses aligned with human preferences. As LLMs become stronger and better…

Computation and Language · Computer Science 2024-06-14 Zhilin Wang , Yi Dong , Olivier Delalleau , Jiaqi Zeng , Gerald Shen , Daniel Egert , Jimmy J. Zhang , Makesh Narsimhan Sreedhar , Oleksii Kuchaiev

By incorporating visual inputs, Multimodal Large Language Models (MLLMs) extend LLMs to support visual reasoning. However, this integration also introduces new vulnerabilities, making MLLMs susceptible to multimodal jailbreak attacks and…

Cryptography and Security · Computer Science 2025-12-04 Beitao Chen , Xinyu Lyu , Lianli Gao , Jingkuan Song , Heng Tao Shen

Activation steering is a promising technique for controlling LLM behavior by adding semantically meaningful vectors directly into a model's hidden states during inference. It is often framed as a precise, interpretable, and potentially…

Machine Learning · Computer Science 2026-02-17 Anton Korznikov , Andrey Galichin , Alexey Dontsov , Oleg Y. Rogov , Ivan Oseledets , Elena Tutubalina

Long context reasoning in large language models (LLMs) has demonstrated enhancement of their cognitive capabilities via chain-of-thought (CoT) inference. Training such models is usually done via reinforcement learning with verifiable…

Computation and Language · Computer Science 2025-12-05 Purbesh Mitra , Sennur Ulukus

Dynamic treatment regimes (DTRs) are critical to precision medicine, optimizing long-term outcomes through personalized, real-time decision-making in evolving clinical contexts, but require careful supervision for unsafe treatment risks.…

Machine Learning · Computer Science 2025-06-10 Yishan Shen , Yuyang Ye , Hui Xiong , Yong Chen

Steering methods have emerged as effective and targeted tools for guiding large language models' (LLMs) behavior without modifying their parameters. Multimodal large language models (MLLMs), however, do not currently enjoy the same suite of…

Machine Learning · Computer Science 2025-05-21 Woody Haosheng Gan , Deqing Fu , Julian Asilis , Ollie Liu , Dani Yogatama , Vatsal Sharan , Robin Jia , Willie Neiswanger

Large Language Models (LLMs) are increasingly used to simulate human users in interactive settings such as therapy, education, and social role-play. While these simulations enable scalable training and evaluation of AI agents, off-the-shelf…

Computation and Language · Computer Science 2025-11-04 Marwa Abdulhai , Ryan Cheng , Donovan Clay , Tim Althoff , Sergey Levine , Natasha Jaques

This paper presents a comprehensive empirical study on the safety alignment capabilities. We evaluate what matters for safety alignment in LLMs and LRMs to provide essential insights for developing more secure and reliable AI systems. We…

Computation and Language · Computer Science 2026-02-25 Xing Li , Hui-Ling Zhen , Lihao Yin , Xianzhi Yu , Zhenhua Dong , Mingxuan Yuan

Recent advancements in language models (LMs) have marked a shift toward the growing importance of post-training. Yet, post-training approaches such as supervised fine-tuning (SFT) do not guarantee the effective use of knowledge acquired…

Computation and Language · Computer Science 2025-10-30 Chunyuan Deng , Ruidi Chang , Hanjie Chen

Large Reasoning Models (LRMs) have advanced automated multi-step reasoning, but their ability to generate complex Chain-of-Thought (CoT) trajectories introduces severe privacy risks, as sensitive information may be deeply embedded…

Artificial Intelligence · Computer Science 2026-01-15 Jingjing Zhou , Gaoxiang Cong , Li Su , Liang Li

Safety is an indispensable requirement for applying reinforcement learning (RL) to real problems. Although there has been a surge of safe RL algorithms proposed in recent years, most existing work typically 1) relies on receiving numeric…

Machine Learning · Computer Science 2024-01-12 Akifumi Wachi , Wataru Hashimoto , Kazumune Hashimoto

Preference-based reinforcement learning (PbRL) bypasses complex reward engineering by learning rewards directly from human preferences, enabling better alignment with human intentions. However, its effectiveness in multi-stage tasks, where…

Machine Learning · Computer Science 2025-09-30 Yao Luan , Ni Mu , Yiqin Yang , Bo Xu , Qing-Shan Jia

Fine-tuning well-aligned large language models (LLMs) on new domains often degrades their safety alignment, even when using benign datasets. Existing safety alignment techniques primarily focus on pretraining, leaving fine-tuned models…

Machine Learning · Computer Science 2026-04-21 Thong Bach , Truyen Tran

Large Language Models (LLMs), despite advances in instruction tuning, often fail to follow complex user instructions. Activation steering techniques aim to mitigate this by manipulating model internals, but have a potential risk of…

Machine Learning · Computer Science 2026-03-10 Minjae Kang , Jaehyung Kim

Safety backdoor attacks in large language models (LLMs) enable the stealthy triggering of unsafe behaviors while evading detection during normal interactions. The high dimensionality of potential triggers in the token space and the diverse…

Cryptography and Security · Computer Science 2024-06-26 Yi Zeng , Weiyu Sun , Tran Ngoc Huynh , Dawn Song , Bo Li , Ruoxi Jia

Steering vectors (SVs) have emerged as a promising approach for interpreting and controlling LLMs, but current methods typically require large contrastive datasets that are often impractical to construct and may capture spurious…

Machine Learning · Computer Science 2025-08-14 Jacob Dunefsky , Arman Cohan
‹ Prev 1 3 4 5 6 7 10 Next ›