English

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents

Sound 2026-08-02 v1

Abstract

We present JoyAI-Talker, a full-duplex speech dialogue system that delivers robust foundation model capabilities while empowering empathetic interaction and voice agent intelligence. JoyAI-Talker adopts a modular Thinker-Talker architecture and further implements a unified speech-text joint training pipeline to mitigate the common "cognitive degradation" bottleneck, thereby largely preserving the model's core textual reasoning, STEM, and logical capabilities while extending them to speech-based interaction. For expressive speech synthesis, the Talker module employs a text-controllable generation paradigm that enables natural-language instructions to flexibly control vocal attributes and localized paralinguistic events, such as laughter and sighs, supporting more expressive and fine-grained speech responses. To enhance conversational empathy, we introduce the Persona-Adaptive Empathetic Response (PAER) framework. PAER employs a hierarchical cognitive pipeline to extract non-verbal speaker cues, such as gender, age, and emotional state, from raw input audio, incorporate them into the Thinker's CoT reasoning, and generate context-adaptive responses that align semantically appropriate text with fine-grained control over utterance-level expressiveness and localized paralinguistic events, including sighs, speaking rate, and volume. We further integrate Joy-Duplex, a state-driven, plug-and-play full-duplex framework that functions as an efficient gating engine for real-time turn control. Extensive evaluations show that JoyAI-Talker achieves highly competitive performance on foundational T2T and S2T benchmarks. In full-duplex evaluation, the system reaches a high response rate of 0.88 under user interruptions while maintaining an extremely low false-trigger rate under background speech, demonstrating its readiness for fluid and natural speech dialogue.

Cite

@article{arxiv.2608.01119,
  title  = {JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents},
  author = {Yinhao Bai and Jinming Chen and Yafeng Chen and Wei Deng and Boya Dong and Nan Duan and Yu Gu and Weisheng Han and Yankun Huang and Ming Ke and Hao Li and Jingdong Li and Xiangyu Liang and Ning Liu and Yuan Liu and Ji Miao and Jiaqi Wang and Qi Wang and Wenchao Wang and Yuxuan Wang and Zhenfang Wang and Zhangyu Xiao and Chao Xue and Hongfei Xue and Fan Yu and Tianyi Zhang and Yuan Zhang and Yuqi Zhang and Lin Zhu},
  journal= {arXiv preprint arXiv:2608.01119},
  year   = {2026}
}