为何为对小模型对个人偏好进行主动人格推断是必要的?
摘要
aligning language models(LMs)到 personalized preferences 的一个突出问题是 underspecification——用户关于 their preferences 的信息不足。一种流行趋势是通过向当前用户的对话添加前缀(例如 prior relevant conversations)来 inject such specification,以 steer preference distribution。大多数方法 passively model personal preferences with prior example preferences pairs。我们询问模型是否从 actively inferring preference descriptions 中受益,并通过创建一个基于 famous people with known public preferences 的 synthetic personalized alignment dataset 来回答此问题。我们 then test how effective finetuned 1-8B size models are at inferring and aligning to personal preferences。结果表明,higher-quality active prefixes 导致 better generalization、more contextually faithful models 以及 less systematic biases across different protected attributes。我们所有结果都表明,active alignment 可以 lead to a more controllable and efficient path for personalized alignment。
引用
@article{arxiv.2505.13257,
title = {Is Active Persona Inference Necessary for Aligning Small Models to Personal Preferences?},
author = {Zilu Tang and Afra Feyza Akyürek and Ekin Akyürek and Derry Wijaya},
journal= {arXiv preprint arXiv:2505.13257},
year = {2025}
}
备注
9 pages, EMNLP PALS workshop 2025