English
Related papers

Related papers: Activation Steering for Bias Mitigation: An Interp…

200 papers

The rapid advancement of large language models (LLMs) has made machine-generated text increasingly difficult to distinguish from human-written text. While recent studies explore leveraging internal representations of language models to…

Applications · Statistics 2026-05-14 Luxu Liang , Xiang Li

As the development and application of Large Language Models (LLMs) continue to advance rapidly, enhancing their trustworthiness and aligning them with human preferences has become a critical area of research. Traditional methods rely…

Computation and Language · Computer Science 2024-11-06 Yuxin Xiao , Chaoqun Wan , Yonggang Zhang , Wenxiao Wang , Binbin Lin , Xiaofei He , Xu Shen , Jieping Ye

Inference-time intervention (ITI) has emerged as a promising method for steering large language model (LLM) behavior in a particular direction (e.g., improving helpfulness) by intervening on token representations without costly updates to…

Computation and Language · Computer Science 2025-07-10 Duy Nguyen , Archiki Prasad , Elias Stengel-Eskin , Mohit Bansal

As large language models (LLMs) become an important way of information access, there have been increasing concerns that LLMs may intensify the spread of unethical content, including implicit bias that hurts certain populations without…

Computation and Language · Computer Science 2025-07-14 Yuchen Wen , Keping Bi , Wei Chen , Jiafeng Guo , Xueqi Cheng

Large language models (LLMs) tend to verbalize confidence scores that are largely detached from their actual accuracy, yet the geometric relationship governing this behavior remain poorly understood. In this work, we present a mechanistic…

Computation and Language · Computer Science 2026-04-02 Miranda Muqing Miao , Lyle Ungar

Large Language Models (LLMs) often exhibit homogenized cultural perspectives. While the World Values Survey (WVS) provides a gold standard for mapping human values, traditional direct prompting of LLMs on WVS often fails to access the…

Computation and Language · Computer Science 2026-05-27 Trung Duc Anh Dang , Sarah Masud

Despite extensive efforts in safety alignment, large language models (LLMs) remain vulnerable to jailbreak attacks. Activation steering offers a training-free defense method but relies on fixed steering coefficients, resulting in suboptimal…

Cryptography and Security · Computer Science 2025-09-22 Weixiang Zhao , Jiahe Guo , Yulin Hu , Yang Deng , An Zhang , Xingyu Sui , Xinyang Han , Yanyan Zhao , Bing Qin , Tat-Seng Chua , Ting Liu

Activation steering methods were shown to be effective in conditioning language model generation by additively intervening over models' intermediate representations. However, the evaluation of these techniques has so far been limited to…

Computation and Language · Computer Science 2024-12-02 Daniel Scalena , Gabriele Sarti , Malvina Nissim

LLMs are deployed globally, yet produce responses biased towards cultures with abundant training data. Existing cultural localization approaches such as prompting or post-training alignment are black-box, hard to control, and do not reveal…

Computation and Language · Computer Science 2026-03-25 Simran Khanuja , Hongbin Liu , Shujian Zhang , John Lambert , Mingqing Chen , Rajiv Mathews , Lun Wang

Large Language Models (LLMs) are powerful tools with the potential to benefit society immensely, yet, they have demonstrated biases that perpetuate societal inequalities. Despite significant advancements in bias mitigation techniques using…

Computation and Language · Computer Science 2024-09-24 Deonna M. Owens , Ryan A. Rossi , Sungchul Kim , Tong Yu , Franck Dernoncourt , Xiang Chen , Ruiyi Zhang , Jiuxiang Gu , Hanieh Deilamsalehy , Nedim Lipka

Existing approaches to bias evaluation in large language models (LLMs) trade ecological validity for statistical control, relying either on artificial prompts that poorly reflect real-world use or on naturalistic tasks that lack scale and…

Computation and Language · Computer Science 2026-05-12 Akram Elbouanani , Aboubacar Tuo , Adrian Popescu

Explainability for Large Language Models (LLMs) is a critical yet challenging aspect of natural language processing. As LLMs are increasingly integral to diverse applications, their "black-box" nature sparks significant concerns regarding…

Computation and Language · Computer Science 2024-02-23 Haoyan Luo , Lucia Specia

We introduce SteeringSafety, a systematic framework for evaluating representation steering methods across seven safety perspectives spanning 17 datasets. While prior work highlights general capabilities of representation steering, we…

Artificial Intelligence · Computer Science 2025-10-17 Vincent Siu , Nicholas Crispino , David Park , Nathan W. Henry , Zhun Wang , Yang Liu , Dawn Song , Chenguang Wang

Most jailbreak techniques for Large Language Models (LLMs) primarily rely on prompt modifications, including paraphrasing, obfuscation, or conversational strategies. Meanwhile, abliteration techniques (also known as targeted ablations of…

Cryptography and Security · Computer Science 2026-03-17 Maël Jenny , Jérémie Dentan , Sonia Vanier , Michaël Krajecki

Adapting models to a language that was only partially present in the pre-training data requires fine-tuning, which is expensive in terms of both data and computational resources. As an alternative to fine-tuning, we explore the potential of…

Computation and Language · Computer Science 2024-11-28 Daniel Scalena , Elisabetta Fersini , Malvina Nissim

Large Language Models (LLMs) like gpt-3.5-turbo-0613 and claude-instant-1.2 are vital in interpreting and executing semantic tasks. Unfortunately, these models' inherent biases adversely affect their performance Particularly affected is…

Computation and Language · Computer Science 2024-06-18 J. E. Eicher , R. F. Irgolič

Transformer-based pretrained large language models (PLM) such as BERT and GPT have achieved remarkable success in NLP tasks. However, PLMs are prone to encoding stereotypical biases. Although a burgeoning literature has emerged on…

Computation and Language · Computer Science 2024-06-18 Yi Yang , Hanyu Duan , Ahmed Abbasi , John P. Lalor , Kar Yan Tam

Given the prompt "Rome is in", can we steer a language model to flip its prediction of an incorrect token "France" to a correct token "Italy" by only multiplying a few relevant activation vectors with scalars? We argue that successfully…

Computation and Language · Computer Science 2024-10-08 Niklas Stoehr , Kevin Du , Vésteinn Snæbjarnarson , Robert West , Ryan Cotterell , Aaron Schein

Large Language Models (LLMs) have transformed the field of artificial intelligence by unlocking the era of generative applications. Built on top of generative AI capabilities, Agentic AI represents a major shift toward autonomous,…

Artificial Intelligence · Computer Science 2025-08-27 Karanbir Singh , Deepak Muppiri , William Ngu

Large language models (LLMs) have shown remarkable success in recent years, enabling a wide range of applications, including intelligent assistants that support users' daily life and work. A critical factor in building such assistants is…

Computation and Language · Computer Science 2025-10-28 Xiaoyan Zhao , Ming Yan , Yilun Qiu , Haoting Ni , Yang Zhang , Fuli Feng , Hong Cheng , Tat-Seng Chua
‹ Prev 1 8 9 10 Next ›