大语言模型的值是否与人类结构对齐?——因果视角
计算与语言
2025-02-25 v2 人工智能
机器学习
摘要
随着大语言模型 (LLMs) 越来越多地集成到关键应用中,对其行为与人类价值观的对齐 presents significant challenges。当前方法,如强化学习从人类反馈 (RLHF),通常 focuses on a limited set of coarse-grained values and are resource-intensive。此外,这些价值之间的相关性 remains implicit,导致对 value-steering outcomes 的解释不清晰。我们的工作认为,大语言模型的价值维度 underlying a latent causal value graph,并且尽管进行对齐训练,这个结构仍显著不同于人类价值体系。我们利用这些因果价值图指导两种轻量级 value-steering 方法:基于角色的提示和稀疏自编码器 (SAE) 驾驶,有效缓解了意外副作用。此外,SAE 提供了更细粒度的价值驾驶方法。Gemma-2B-IT 和 Llama3-8B-IT 上的实验表明,我们的方法在有效性和可控性方面都表现出色。
引用
@article{arxiv.2501.00581,
title = {Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective},
author = {Yipeng Kang and Junqi Wang and Yexin Li and Mengmeng Wang and Wenming Tu and Quansen Wang and Hengli Li and Tingjun Wu and Xue Feng and Fangwei Zhong and Zilong Zheng},
journal= {arXiv preprint arXiv:2501.00581},
year = {2025}
}