中文

基于 LLM 人格模拟的数字孪生体:TwinVoice 基准

计算与语言 2025-10-31 v2

摘要

大型语言模型 (Large Language Models, LLMs) 正显示出人类水平的能力,越来越被视为模拟个体沟通风格、行为倾向和性格特征的基础。然而,现有的 LLM 人格模拟评估仍受限:大多数依赖合成对话,缺乏系统性框架,且缺乏对能力要求的分析。为此,我们引入 TwinVoice,这是一个用于评估人格模拟在多样化真实情境下的表现的全面基准测试。TwinVoice 包含三个维度:社交人格 (Social Persona)、人际人格 (Interpersonal Persona) 以及叙事人格 (Narrative Persona)。进一步将 LLM 性能评估分解为六项基本能力,包括意见一致性、记忆回顾、逻辑推理、词汇保真度、人格语调以及句法风格。实验结果显示,虽然先进模型在人格模拟方面实现了中等准确度,但在诸如句法风格和记忆回顾等能力方面仍不足。因此,LLM 实现的平均性能远低于人类基准。

关键词

引用

@article{arxiv.2510.25536,
  title  = {TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation},
  author = {Bangde Du and Minghao Guo and Songming He and Ziyi Ye and Xi Zhu and Weihang Su and Shuqi Zhu and Yujia Zhou and Yongfeng Zhang and Qingyao Ai and Yiqun Liu},
  journal= {arXiv preprint arXiv:2510.25536},
  year   = {2025}
}

备注

Main paper: 11 pages, 3 figures, 6 tables. Appendix: 28 pages. Bangde Du and Minghao Guo contributed equally. Corresponding authors: Ziyi Ye ([email protected]), Qingyao Ai ([email protected])