中文
相关论文

相关论文: Activation-Space Personality Steering: Hybrid Laye…

200 篇论文

Large language models (LLMs) require precise behavior control for safe and effective deployment across diverse applications. Activation steering offers a promising approach for LLMs' behavioral control. We focus on the question of how…

人工智能 · 计算机科学 2026-01-13 Tetiana Bas , Krystian Novak

Using psychological constructs such as the Big Five, large language models (LLMs) can imitate specific personality profiles and predict a user's personality. While LLMs can exhibit behaviors consistent with these constructs, it remains…

计算与语言 · 计算机科学 2026-04-14 Yuto Harada , Hiro Taiyo Hamada

As Large Language Models (LLMs) become integral to human-centered applications, understanding their personality-like behaviors is increasingly important for responsible development and deployment. This paper systematically evaluates six…

计算与语言 · 计算机科学 2025-11-07 Christos-Nikolaos Zacharopoulos , Revekka Kyriakoglou

We present Big5-Scaler, a prompt-based framework for conditioning large language models (LLMs) with controllable Big Five personality traits. By embedding numeric trait values into natural language prompts, our method enables fine-grained…

计算与语言 · 计算机科学 2025-08-11 Gunhee Cho , Yun-Gyung Cheong

Inference-time LLM alignment methods, particularly activation steering, offer an alternative to fine-tuning by directly modifying activations during generation. Existing methods, however, often rely on non-anticipative interventions that…

机器学习 · 计算机科学 2026-04-22 Julian Skifstad , Xinyue Annie Yang , Glen Chou

Large Language Models (LLMs) have demonstrated human-like capabilities in language comprehension and generation, becoming active participants in social and cognitive domains. This study investigates whether LLMs exhibit personality-like…

计算与语言 · 计算机科学 2025-05-22 Wang Jiaqi , Wang bo , Guo fa , Cheng cheng , Yang li

Recent mechanistic studies suggest that large language models (LLMs) may utilize their depth inefficiently in standard single-turn tasks. Whether this still holds in autonomous agent settings, where models must perform multi-turn planning,…

人工智能 · 计算机科学 2026-05-28 Zhenyu Cui , Xiangzhong Luo

Changing the behavior of large language models (LLMs) can be as straightforward as editing the Transformer's residual streams using appropriately constructed "steering vectors." These modifications to internal neural activations, a form of…

计算与语言 · 计算机科学 2025-05-20 Jian-Qiao Zhu , Haijiang Yan , Thomas L. Griffiths

Large Language Models (LLMs) have demonstrated promising capabilities to generate responses that simulate consistent personality traits. Despite the major attempts to analyze personality expression through output-based evaluations, little…

计算与语言 · 计算机科学 2025-07-30 Tianjie Ju , Zhenyu Shao , Bowen Wang , Yujia Chen , Zhuosheng Zhang , Hao Fei , Mong-Li Lee , Wynne Hsu , Sufeng Duan , Gongshen Liu

Recent advancements in Large Language Models (LLMs) have led to their adaptation in various domains as conversational agents. We wonder: can personality tests be applied to these agents to analyze their behavior, similar to humans? We…

Large Language Models (LLMs) are increasingly deployed as autonomous agents, necessitating a deeper understanding of their decision-making behaviour under risk. This study investigates the relationship between LLMs' personality traits and…

计算机与社会 · 计算机科学 2025-03-10 John Hartley , Conor Hamill , Devesh Batra , Dale Seddon , Ramin Okhrati , Raad Khraishi

The emergence of unveiling human-like behaviors in Large Language Models (LLMs) has led to a closer connection between NLP and human psychology. Scholars have been studying the inherent personalities exhibited by LLMs and attempting to…

人工智能 · 计算机科学 2025-03-25 Lucio La Cava , Andrea Tagarelli

This research explores strategies for steering the output of large language models (LLMs) towards specific styles, such as sentiment, emotion, or writing style, by adding style vectors to the activations of hidden layers during text…

Recent work shows that large language models (LLMs) encode behavioural traits ("personas") as linear directions in activation space, often called "persona vectors". Prior work has used such directions as static handles for behavioural…

人工智能 · 计算机科学 2026-05-12 Nils A. Herrmann , Leander Girrbach , Kirill Bykov , Zeynep Akata

This study validates Large Language Models (LLMs) as a dynamic alternative to questionnaire-based personality assessment. Using a within-subjects experiment (N=33), we compared Big Five personality scores derived from guided LLM…

计算与语言 · 计算机科学 2026-02-19 Andrius Matšenas , Anet Lello , Tõnis Lees , Hans Peep , Kim Lilii Tamm

Steering vectors have emerged as a lightweight and effective approach for aligning large language models (LLMs) at inference time, enabling modulation over model behaviors by shifting LLM representations towards a target behavior. However,…

机器学习 · 计算机科学 2026-04-07 Soham Gadgil , Chris Lin , Su-In Lee

Large language models (LLMs) are increasingly deployed as autonomous decision-makers in strategic settings, yet we have limited tools for understanding their high-level behavioral traits. We use activation steering methods in game-theoretic…

人工智能 · 计算机科学 2026-03-24 Johnathan Sun , Andrew Zhang

Large language models (LLMs) excel in both closed tasks (including problem-solving, and code generation) and open tasks (including creative writing), yet existing explanations for their capabilities lack connections to real-world human…

计算与语言 · 计算机科学 2025-05-28 Yifan Duan , Yihong Tang , Xuefeng Bai , Kehai Chen , Juntao Li , Min Zhang

Recent research has explored LLMs as scalable tools for relevance labeling, but studies indicate they are susceptible to priming effects, where prior relevance judgments influence later ones. Although psychological theories link personality…

计算与语言 · 计算机科学 2025-12-02 Nuo Chen , Hanpei Fang , Jiqun Liu , Wilson Wei , Tetsuya Sakai , Xiao-Ming Wu

Large language models (LLMs) enable conversational agents (CAs) to express distinctive personalities, raising new questions about how such designs shape user perceptions. This study investigates how personality expression levels and…

人机交互 · 计算机科学 2026-04-30 Hasibur Rahman , Smit Desai