English
Related papers

Related papers: Selective Steering: Norm-Preserving Control Throug…

200 papers

Activation steering methods are widely used to control large language model (LLM) behavior and are often interpreted as revealing meaningful internal representations. This interpretation assumes that steering directions are identifiable and…

Machine Learning · Computer Science 2026-04-02 Sohan Venkatesh , Ashish Mahendran Kurapath

Large Language Models (LLMs) are increasingly deployed in high-stakes decision-making contexts. While prior work has shown that LLMs exhibit cognitive biases behaviorally, whether these biases correspond to identifiable internal…

Artificial Intelligence · Computer Science 2026-04-03 Fan Huang , Songheng Zhang , Haewoon Kwak , Jisun An

Ensuring robust safety measures across a wide range of scenarios is crucial for user-facing systems. While Large Language Models (LLMs) can generate valuable data for safety measures, they often exhibit distributional biases, focusing on…

Computation and Language · Computer Science 2024-10-16 Sabit Hassan , Anthony Sicilia , Malihe Alikhani

Methods for controlling large language models (LLMs), including local weight fine-tuning, LoRA-based adaptation, and activation-based interventions, are often studied in isolation, obscuring their connections and making comparison…

Computation and Language · Computer Science 2026-04-14 Ziwen Xu , Chenyan Wu , Hengyu Sun , Haiwen Hong , Mengru Wang , Yunzhi Yao , Longtao Huang , Hui Xue , Shumin Deng , Zhixuan Chu , Huajun Chen , Ningyu Zhang

The growing use of generative models in daily life calls for efficient mechanisms to control their generation, to e.g., produce safe content or provide users with tools to explore style changes. Ideally, such mechanisms should require low…

Computation and Language · Computer Science 2025-10-20 Pau Rodriguez , Michal Klein , Eleonora Gualdoni , Valentino Maiorca , Arno Blaas , Luca Zappella , Marco Cuturi , Xavier Suau

Modern large language models (LLMs) are typically secured by auditing data, prompts, and refusal policies, while treating the forward pass as an implementation detail. We show that intermediate activations in decoder-only LLMs form a…

Cryptography and Security · Computer Science 2025-11-24 Zhiyuan Xu , Stanislav Abaimov , Joseph Gardiner , Sana Belguith

Steering language model generation toward desired textual properties is essential for practical deployment, and inference-time methods are particularly appealing because they enable controllable generation without retraining. Recent work…

Computation and Language · Computer Science 2026-05-29 Hyeseon An , Yo-Sub Han

While the wide adoption of refusal training in large language models (LLMs) has showcased improvements in model safety, recent works have highlighted shortcomings due to the shallow nature of these alignment methods. To this end, the work…

Machine Learning · Computer Science 2026-04-17 Pankayaraj Pathmanathan , Furong Huang

Large Audio-Language Models (LALMs) are becoming essential as a powerful multimodal backbone for real-world applications. However, recent studies show that audio inputs can more easily elicit harmful responses than text, exposing new risks…

Sound · Computer Science 2026-05-08 Weilin Lin , Jianze Li , Hui Xiong , Li Liu

Understanding and controlling the behavior of large language models (LLMs) is an increasingly important topic in multilingual NLP. Beyond prompting or fine-tuning, , i.e.,~manipulating internal representations during inference, has emerged…

Steering large language models (LLMs) is usually done by either instruction prompting or activation steering. Prompting often gives strong control, but caches guidance tokens at every layer and can clutter long interactions; activation…

Machine Learning · Computer Science 2026-05-12 Andy Zeyi Liu , Michael Zhang , Ilana Greenberg , Adam Alnasser , Lucas Baker , John Sous

Preference alignment is a critical step in making Large Language Models (LLMs) useful and aligned with (human) preferences. Existing approaches such as Reinforcement Learning from Human Feedback or Direct Preference Optimization typically…

Computation and Language · Computer Science 2025-09-30 Lucio La Cava , Andrea Tagarelli

Test-time compute has emerged as a powerful paradigm for improving the performance of large language models (LLMs), where generating multiple outputs or refining individual chains can significantly boost answer accuracy. However, existing…

Machine Learning · Computer Science 2025-09-26 Sheng Liu , Tianlang Chen , Pan Lu , Haotian Ye , Yizheng Chen , Lei Xing , James Zou

Large Language Models (LLMs) hold immense potential to generate synthetic data of high quality and utility, which has numerous applications from downstream model training to practical data utilisation. However, contemporary models, despite…

Computation and Language · Computer Science 2023-08-21 Charles O'Neill , Yuan-Sen Ting , Ioana Ciuca , Jack Miller , Thang Bui

Large language models have transformed AI, yet reliably controlling their outputs remains a challenge. This paper explores activation engineering, where outputs of pre-trained LLMs are controlled by manipulating their activations at…

Neural and Evolutionary Computing · Computer Science 2025-05-13 Joris Postmus , Steven Abreu

Inspired by the exceptional general intelligence of Large Language Models (LLMs), researchers have begun to explore their application in pioneering the next generation of recommender systems - systems that are conversational, explainable,…

Information Retrieval · Computer Science 2024-08-06 Wensheng Lu , Jianxun Lian , Wei Zhang , Guanghua Li , Mingyang Zhou , Hao Liao , Xing Xie

Activation steering, or representation engineering, offers a lightweight approach to align large language models (LLMs) by manipulating their internal activations at inference time. However, current methods suffer from two key limitations:…

Artificial Intelligence · Computer Science 2026-02-24 Hongjue Zhao , Haosen Sun , Jiangtao Kong , Xiaochang Li , Qineng Wang , Liwei Jiang , Qi Zhu , Tarek Abdelzaher , Yejin Choi , Manling Li , Huajie Shao

With the growing adoption of Large Language Models (LLMs) in critical areas, ensuring their security against jailbreaking attacks is paramount. While traditional defenses primarily rely on refusing malicious prompts, recent logit-level…

Cryptography and Security · Computer Science 2025-07-31 Yassine Rachidy , Jihad Rbaiti , Youssef Hmamouche , Faissal Sehbaoui , Amal El Fallah Seghrouchni

The rise of large language models (LLMs) has prompted increasing interest in their use as in-context learning agents. At the core of agentic behavior is the capacity for exploration, or the ability to actively gather information about the…

Computation and Language · Computer Science 2024-10-14 Nate Rahn , Pierluca D'Oro , Marc G. Bellemare

Large language models (LLMs) emulate a consistent human-like behavior that can be shaped through activation-level interventions. This paradigm is converging on additive residual-stream injections, which rely on injection-strength sweeps to…

Computation and Language · Computer Science 2026-04-17 Leonardo Blas , Robin Jia , Emilio Ferrara
‹ Prev 1 4 5 6 7 8 10 Next ›