English

DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation

Sound 2025-10-02 v2 Audio and Speech Processing

Abstract

Neural audio codecs form the foundational building blocks for language model (LM)-based speech generation. Typically, there is a trade-off between frame rate and audio quality. This study introduces a low-frame-rate, semantically enhanced codec model. Existing approaches distill semantically rich self-supervised (SSL) representations into the first-layer codec tokens. This work proposes DualCodec, a dual-stream encoding approach that integrates SSL and waveform representations within an end-to-end codec framework. In this setting, DualCodec enhances the semantic information in the first-layer codec and enables the codec system to maintain high audio quality while operating at a low frame rate. Note that a low-frame-rate codec improves the efficiency of speech generation. Experimental results on audio codec and speech generation tasks confirm the effectiveness of the proposed DualCodec compared to state-of-the-art codec systems, such as Mimi Codec, SpeechTokenizer, DAC, and Encodec. Demos are available at: https://dualcodec.github.io, code is available at: https://github.com/jiaqili3/DualCodec

Keywords

Cite

@article{arxiv.2505.13000,
  title  = {DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation},
  author = {Jiaqi Li and Xiaolong Lin and Zhekai Li and Shixi Huang and Yuancheng Wang and Chaoren Wang and Zhenpeng Zhan and Zhizheng Wu},
  journal= {arXiv preprint arXiv:2505.13000},
  year   = {2025}
}

Comments

Accepted to Interspeech 2025

R2 v1 2026-07-01T02:21:35.871Z