English
Related papers

Related papers: Extending Activation Steering to Broad Skills and …

200 papers

Prior work argues that refusal in large language models is mediated by a single activation-space direction, enabling effective steering and ablation. We show that this account is incomplete. Across eleven categories of refusal and…

Computation and Language · Computer Science 2026-02-03 Faaiz Joad , Majd Hawasly , Sabri Boughorbel , Nadir Durrani , Husrev Taha Sencar

The mechanisms by which reasoning training reshapes LLMs' internal computations remain unclear. We study lightweight steering vectors inserted into the base model's residual stream and trained with a reinforcement-learning objective. These…

With the emergence of large language models (LLMs) as a powerful class of generative artificial intelligence (AI), their use in tutoring has become increasingly prominent. Prior works on LLM-based tutoring typically learn a single tutor…

Computation and Language · Computer Science 2026-02-10 Jaewook Lee , Alexander Scarlatos , Simon Woodhead , Andrew Lan

Driving requires reacting to a wide variety of complex environment conditions and agent behaviors. Explicitly modeling each possible scenario is unrealistic. In contrast, imitation learning can, in theory, leverage data from large fleets of…

Computer Vision and Pattern Recognition · Computer Science 2019-04-22 Felipe Codevilla , Eder Santana , Antonio M. López , Adrien Gaidon

Prompt engineering and finetuning aim to maximize language model performance on a given metric (like toxicity reduction). However, these methods do not fully elicit a model's capabilities. To reduce this gap, we introduce activation…

Computation and Language · Computer Science 2024-10-11 Alexander Matt Turner , Lisa Thiergart , Gavin Leech , David Udell , Juan J. Vazquez , Ulisse Mini , Monte MacDiarmid

Activation steering has emerged as a cost-effective paradigm for modifying large language model (LLM) behaviors. Existing methods typically intervene at the block level, steering the bundled activations of selected attention heads,…

Computation and Language · Computer Science 2026-02-05 Zijian Feng , Tianjiao Li , Zixiao Zhu , Hanzhang Zhou , Junlang Qian , Li Zhang , Jia Jim Deryl Chua , Lee Onn Mak , Gee Wah Ng , Kezhi Mao

Large Language Models (LLMs) often generate inconsistent responses when prompted with semantically equivalent paraphrased inputs. Recently, activation steering, a technique that modulates LLMs' behaviours by adjusting their latent…

Computation and Language · Computer Science 2025-01-23 Jingyuan Yang , Rongjun Li , Weixuan Wang , Ziyu Zhou , Zhiyong Feng , Wei Peng

Activation steering is a widely used approach for controlling large language model (LLM) behavior by intervening on internal representations. Existing methods largely rely on the Linear Representation Hypothesis, assuming behavioral…

Artificial Intelligence · Computer Science 2026-03-24 Shivam Raval , Hae Jin Song , Linlin Wu , Abir Harrasse , Jeff M. Phillips , Fazl Barez , Amirali Abdullah

We show that iterative deployment of large language models (LLMs), each fine-tuned on data carefully curated by users from the previous models' deployment, can significantly change the properties of the resultant models. By testing this…

Artificial Intelligence · Computer Science 2026-01-01 Augusto B. Corrêa , Yoav Gelberg , Luckeciano C. Melo , Ilia Shumailov , André G. Pereira , Yarin Gal

Large Language Models (LLMs) have shown impressive capabilities across a wide variety of tasks. However, they still face challenges with long-horizon planning. To study this, we propose path planning tasks as a platform to evaluate LLMs'…

Artificial Intelligence · Computer Science 2024-06-24 Mohamed Aghzal , Erion Plaku , Ziyu Yao

Imitation learning has demonstrated strong performance in robotic manipulation by learning from large-scale human demonstrations. While existing models excel at single-task learning, it is observed in practical applications that their…

Robotics · Computer Science 2026-01-21 Wangtian Shen , Jinming Ma , Mingliang Zhou , Ziyang Meng

Autonomous robots combine a variety of skills to form increasingly complex behaviors called missions. While the skills are often programmed at a relatively low level of abstraction, their coordination is architecturally separated and often…

Robotics · Computer Science 2020-11-17 Razan Ghzouli , Thorsten Berger , Einar Broch Johnsen , Swaib Dragule , Andrzej Wąsowski

Imitation learning has driven the development of generalist policies capable of autonomously solving multiple tasks. However, when a pretrained policy makes errors during deployment, there are limited mechanisms for users to correct its…

Robotics · Computer Science 2025-06-18 Yanwei Wang

Language models (LMs) are increasingly used in high-stakes, multi-agent settings, where following instructions and maintaining value alignment are critical. Most alignment research focuses on interactions between a single LM and a single…

Artificial Intelligence · Computer Science 2026-05-12 Maria Chang , Ronny Luss , Miao Liu , Keerthiram Murugesan , Karthikeyan Ramamurthy , Djallel Bouneffouf

Personality steering in large language models (LLMs) commonly relies on injecting trait-specific steering vectors, implicitly assuming that personality traits can be controlled independently. In this work, we examine whether this assumption…

Computation and Language · Computer Science 2026-02-19 Pranav Bhandari , Usman Naseem , Mehwish Nasim

Modern large language models (LLMs) are typically secured by auditing data, prompts, and refusal policies, while treating the forward pass as an implementation detail. We show that intermediate activations in decoder-only LLMs form a…

Cryptography and Security · Computer Science 2025-11-24 Zhiyuan Xu , Stanislav Abaimov , Joseph Gardiner , Sana Belguith

The sparsely-activated models have achieved great success in natural language processing through large-scale parameters and relatively low computational cost, and gradually become a feasible technique for training and implementing extremely…

Activation steering has emerged as a promising approach for efficiently adapting large language models (LLMs) to downstream behaviors. However, most existing steering methods rely on a single static direction per task or concept, making…

Feature steering has emerged as a promising approach for controlling LLM behavior through direct manipulation of internal representations, offering advantages over prompt engineering. However, its practical effectiveness in real-world…

Advanced driver assistance systems have successfully reduced drivers' workloads and increased safety. On the other hand, the excessive use of such systems can impede the development of driving skills. However, there exist collaborative…

Human-Computer Interaction · Computer Science 2019-07-02 Takahiro Wada
‹ Prev 1 3 4 5 6 7 10 Next ›