Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation

Hanzhao Li; Liumeng Xue; Haohan Guo; Xinfa Zhu; Yuanjun Lv; Lei Xie; Yunlin Chen; Hao Yin; Zhifei Li

Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation

Audio and Speech Processing 2024-06-12 v1

Authors: Hanzhao Li , Liumeng Xue , Haohan Guo , Xinfa Zhu , Yuanjun Lv , Lei Xie , Yunlin Chen , Hao Yin , Zhifei Li

Abstract

The multi-codebook speech codec enables the application of large language models (LLM) in TTS but bottlenecks efficiency and robustness due to multi-sequence prediction. To avoid this obstacle, we propose Single-Codec, a single-codebook single-sequence codec, which employs a disentangled VQ-VAE to decouple speech into a time-invariant embedding and a phonetically-rich discrete sequence. Furthermore, the encoder is enhanced with 1) contextual modeling with a BLSTM module to exploit the temporal information, 2) a hybrid sampling module to alleviate distortion from upsampling and downsampling, and 3) a resampling module to encourage discrete units to carry more phonetic information. Compared with multi-codebook codecs, e.g., EnCodec and TiCodec, Single-Codec demonstrates higher reconstruction quality with a lower bandwidth of only 304bps. The effectiveness of Single-Code is further validated by LLM-TTS experiments, showing improved naturalness and intelligibility.

Keywords

speech processing speech recognition and language modeling

Cite

@article{arxiv.2406.07422,
  title  = {Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation},
  author = {Hanzhao Li and Liumeng Xue and Haohan Guo and Xinfa Zhu and Yuanjun Lv and Lei Xie and Yunlin Chen and Hao Yin and Zhifei Li},
  journal= {arXiv preprint arXiv:2406.07422},
  year   = {2024}
}

Comments

Accepted by Interspeech 2024

Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation

Abstract

Keywords

Cite

Comments

Related papers