English

Building a Taiwanese Mandarin Spoken Language Model: A First Attempt

Computation and Language 2024-12-30 v2 Sound Audio and Speech Processing

Abstract

This technical report presents our initial attempt to build a spoken large language model (LLM) for Taiwanese Mandarin, specifically tailored to enable real-time, speech-to-speech interaction in multi-turn conversations. Our end-to-end model incorporates a decoder-only transformer architecture and aims to achieve seamless interaction while preserving the conversational flow, including full-duplex capabilities allowing simultaneous speaking and listening. The paper also details the training process, including data preparation with synthesized dialogues and adjustments for real-time interaction. We also developed a platform to evaluate conversational fluency and response coherence in multi-turn dialogues. We hope the release of the report can contribute to the future development of spoken LLMs in Taiwanese Mandarin.

Keywords

Cite

@article{arxiv.2411.07111,
  title  = {Building a Taiwanese Mandarin Spoken Language Model: A First Attempt},
  author = {Chih-Kai Yang and Yu-Kuan Fu and Chen-An Li and Yi-Cheng Lin and Yu-Xiang Lin and Wei-Chih Chen and Ho Lam Chung and Chun-Yi Kuan and Wei-Ping Huang and Ke-Han Lu and Tzu-Quan Lin and Hsiu-Hsuan Wang and En-Pei Hu and Chan-Jan Hsu and Liang-Hsuan Tseng and I-Hsiang Chiu and Ulin Sanga and Xuanjun Chen and Po-chun Hsu and Shu-wen Yang and Hung-yi Lee},
  journal= {arXiv preprint arXiv:2411.07111},
  year   = {2024}
}

Comments

Work in progress

R2 v1 2026-06-28T19:55:44.916Z