中文
相关论文

相关论文: Persona Vectors: Monitoring and Controlling Charac…

200 篇论文

Large language models (LLMs) are increasingly deployed as autonomous decision-makers in strategic settings, yet we have limited tools for understanding their high-level behavioral traits. We use activation steering methods in game-theoretic…

人工智能 · 计算机科学 2026-03-24 Johnathan Sun , Andrew Zhang

Large language models can represent a variety of personas but typically default to a helpful Assistant identity cultivated during post-training. We investigate the structure of the space of model personas by extracting activation directions…

计算与语言 · 计算机科学 2026-01-16 Christina Lu , Jack Gallagher , Jonathan Michala , Kyle Fish , Jack Lindsey

Driven by the demand for personalized AI systems, there is growing interest in aligning the behavior of large language models (LLMs) with human traits such as personality. Previous attempts to induce personality in LLMs have shown promising…

计算与语言 · 计算机科学 2025-09-25 Seungjong Sun , Seo Yeon Baek , Jang Hyun Kim

How large language models internally represent high-level behaviors is a core interpretability question with direct relevance to AI safety: it determines what we can detect, audit, or intervene on. Recent work has shown that traits such as…

计算与语言 · 计算机科学 2026-05-14 Viktor Moskvoretskii , Dominik Glandorf , Jorge Medina Moreira , Tanja Käser , Robert West

It is becoming increasingly necessary to have monitors check for harmful behaviors during language model interactions, but text-only monitoring has not been sufficient. This is because models sometimes exhibit strategic deception and…

人工智能 · 计算机科学 2026-05-18 Prasad Mahadik , Adrians Skapars

Current methods for personality control in Large Language Models rely on static prompting or expensive fine-tuning, failing to capture the dynamic and compositional nature of human traits. We introduce PERSONA, a training-free framework…

人工智能 · 计算机科学 2026-02-18 Xiachong Feng , Liang Zhao , Weihong Zhong , Yichong Huang , Yuxuan Gu , Lingpeng Kong , Xiaocheng Feng , Bing Qin

One way to personalize and steer generations from large language models (LLM) is to assign a persona: a role that describes how the user expects the LLM to behave (e.g., a helpful assistant, a teacher, a woman). This paper investigates how…

计算与语言 · 计算机科学 2025-07-02 Pedro Henrique Luz de Araujo , Benjamin Roth

The influence of personas on Large Language Models (LLMs) has been widely studied, yet their direct impact on performance remains uncertain. This work explores a novel approach to guiding LLM behaviour through role vectors, an alternative…

计算与语言 · 计算机科学 2025-02-18 Daniele Potertì , Andrea Seveso , Fabio Mercorio

We study the effect of different persona on \textbf{sycophancy}: model's agreement with users even when the user is incorrect. The standard mitigation, Contrastive Activation Addition (CAA), derives a steering direction from labelled pairs…

人工智能 · 计算机科学 2026-05-21 Ishaan Kelkar , Nebras Alam , Vikram Kakaria , Madhur Panwar , Vasu Sharma , Maheep Chaudhary

With the emergence of large language models (LLMs) as a powerful class of generative artificial intelligence (AI), their use in tutoring has become increasingly prominent. Prior works on LLM-based tutoring typically learn a single tutor…

计算与语言 · 计算机科学 2026-02-10 Jaewook Lee , Alexander Scarlatos , Simon Woodhead , Andrew Lan

Large language models (LLMs) can be said to have preferences: they reliably pick certain tasks and outputs over others, and preferences shaped by post-training and system prompts appear to shape much of their behaviour. But models can also…

计算与语言 · 计算机科学 2026-05-19 Oscar Gilg , Pierre Beckmann , Daniel Paleka , Patrick Butlin

Researchers have been studying approaches to steer the behavior of Large Language Models (LLMs) and build personalized LLMs tailored for various applications. While fine-tuning seems to be a direct solution, it requires substantial…

计算与语言 · 计算机科学 2024-07-31 Yuanpu Cao , Tianrong Zhang , Bochuan Cao , Ziyi Yin , Lu Lin , Fenglong Ma , Jinghui Chen

Human feedback is commonly utilized to finetune AI assistants. But human feedback may also encourage model responses that match user beliefs over truthful ones, a behaviour known as sycophancy. We investigate the prevalence of sycophancy in…

As language models continue to scale in size and capability, they display an array of emerging behaviors, both beneficial and concerning. This heightens the need to control model behaviors. We hope to be able to control the personality…

计算与语言 · 计算机科学 2024-02-16 Yixuan Weng , Shizhu He , Kang Liu , Shengping Liu , Jun Zhao

Recent work shows that large language models (LLMs) encode behavioural traits ("personas") as linear directions in activation space, often called "persona vectors". Prior work has used such directions as static handles for behavioural…

人工智能 · 计算机科学 2026-05-12 Nils A. Herrmann , Leander Girrbach , Kirill Bykov , Zeynep Akata

Psychology research has long explored aspects of human personality such as extroversion, agreeableness and emotional stability. Categorizations like the `Big Five' personality traits are commonly used to assess and diagnose personality…

人工智能 · 计算机科学 2022-12-21 Graham Caron , Shashank Srivastava

Activation-based steering can personalize large language models at inference time, but its effects in educational settings remain unclear. We study persona vectors for seven character traits in short-answer generation and automated scoring…

计算与语言 · 计算机科学 2026-04-09 Yongchao Wu , Aron Henriksson

Procedural content generation has enabled vast virtual worlds through levels, maps, and quests, but large-scale character generation remains underexplored. We identify two alignment-induced biases in existing methods: a positive moral bias,…

计算与语言 · 计算机科学 2026-05-05 Maan Qraitem , Kate Saenko , Bryan A. Plummer

Personality traits have long been studied as predictors of human behavior. Recent advances in Large Language Models (LLMs) suggest similar patterns may emerge in artificial systems, with advanced LLMs displaying consistent behavioral…

Large language models (LLMs) have shown remarkable promise in simulating human language and behavior. This study investigates how integrating persona variables-demographic, social, and behavioral factors-impacts LLMs' ability to simulate…

计算与语言 · 计算机科学 2024-06-18 Tiancheng Hu , Nigel Collier
‹ 上一页 1 2 3 10 下一页 ›