English

TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation

Computation and Language 2025-10-31 v2

Abstract

Large Language Models (LLMs) are exhibiting emergent human-like abilities and are increasingly envisioned as the foundation for simulating an individual's communication style, behavioral tendencies, and personality traits. However, current evaluations of LLM-based persona simulation remain limited: most rely on synthetic dialogues, lack systematic frameworks, and lack analysis of the capability requirement. To address these limitations, we introduce TwinVoice, a comprehensive benchmark for assessing persona simulation across diverse real-world contexts. TwinVoice encompasses three dimensions: Social Persona (public social interactions), Interpersonal Persona (private dialogues), and Narrative Persona (role-based expression). It further decomposes the evaluation of LLM performance into six fundamental capabilities, including opinion consistency, memory recall, logical reasoning, lexical fidelity, persona tone, and syntactic style. Experimental results reveal that while advanced models achieve moderate accuracy in persona simulation, they still fall short of capabilities such as syntactic style and memory recall. Consequently, the average performance achieved by LLMs remains considerably below the human baseline.

Keywords

Cite

@article{arxiv.2510.25536,
  title  = {TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation},
  author = {Bangde Du and Minghao Guo and Songming He and Ziyi Ye and Xi Zhu and Weihang Su and Shuqi Zhu and Yujia Zhou and Yongfeng Zhang and Qingyao Ai and Yiqun Liu},
  journal= {arXiv preprint arXiv:2510.25536},
  year   = {2025}
}

Comments

Main paper: 11 pages, 3 figures, 6 tables. Appendix: 28 pages. Bangde Du and Minghao Guo contributed equally. Corresponding authors: Ziyi Ye (ziyiye@fudan.edu.cn), Qingyao Ai (aiqy@tsinghua.edu.cn)

R2 v1 2026-07-01T07:11:55.515Z