English

PolyVoice: Language Models for Speech to Speech Translation

Computation and Language 2023-06-14 v2 Audio and Speech Processing

Abstract

We propose PolyVoice, a language model-based framework for speech-to-speech translation (S2ST) system. Our framework consists of two language models: a translation language model and a speech synthesis language model. We use discretized speech units, which are generated in a fully unsupervised way, and thus our framework can be used for unwritten languages. For the speech synthesis part, we adopt the existing VALL-E X approach and build a unit-based audio language model. This grants our framework the ability to preserve the voice characteristics and the speaking style of the original speech. We examine our system on Chinese \rightarrow English and English \rightarrow Spanish pairs. Experimental results show that our system can generate speech with high translation quality and audio quality. Speech samples are available at https://speechtranslation.github.io/polyvoice.

Keywords

Cite

@article{arxiv.2306.02982,
  title  = {PolyVoice: Language Models for Speech to Speech Translation},
  author = {Qianqian Dong and Zhiying Huang and Qiao Tian and Chen Xu and Tom Ko and Yunlong Zhao and Siyuan Feng and Tang Li and Kexin Wang and Xuxin Cheng and Fengpeng Yue and Ye Bai and Xi Chen and Lu Lu and Zejun Ma and Yuping Wang and Mingxuan Wang and Yuxuan Wang},
  journal= {arXiv preprint arXiv:2306.02982},
  year   = {2023}
}
R2 v1 2026-06-28T10:56:48.056Z