English

Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis

Sound 2025-05-20 v1 Audio and Speech Processing

Abstract

Conversational Speech Synthesis (CSS) aims to align synthesized speech with the emotional and stylistic context of user-agent interactions to achieve empathy. Current generative CSS models face interpretability limitations due to insufficient emotional perception and redundant discrete speech coding. To address the above issues, we present Chain-Talker, a three-stage framework mimicking human cognition: Emotion Understanding derives context-aware emotion descriptors from dialogue history; Semantic Understanding generates compact semantic codes via serialized prediction; and Empathetic Rendering synthesizes expressive speech by integrating both components. To support emotion modeling, we develop CSS-EmCap, an LLM-driven automated pipeline for generating precise conversational speech emotion captions. Experiments on three benchmark datasets demonstrate that Chain-Talker produces more expressive and empathetic speech than existing methods, with CSS-EmCap contributing to reliable emotion modeling. The code and demos are available at: https://github.com/AI-S2-Lab/Chain-Talker.

Keywords

Cite

@article{arxiv.2505.12597,
  title  = {Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis},
  author = {Yifan Hu and Rui Liu and Yi Ren and Xiang Yin and Haizhou Li},
  journal= {arXiv preprint arXiv:2505.12597},
  year   = {2025}
}

Comments

16 pages, 5 figures, 5 tables. Accepted by ACL 2025 (Findings)