English

UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units

Computation and Language 2023-05-29 v2 Sound Audio and Speech Processing

Abstract

Direct speech-to-speech translation (S2ST), in which all components can be optimized jointly, is advantageous over cascaded approaches to achieve fast inference with a simplified pipeline. We present a novel two-pass direct S2ST architecture, UnitY, which first generates textual representations and predicts discrete acoustic units subsequently. We enhance the model performance by subword prediction in the first-pass decoder, advanced two-pass decoder architecture design and search strategy, and better training regularization. To leverage large amounts of unlabeled text data, we pre-train the first-pass text decoder based on the self-supervised denoising auto-encoding task. Experimental evaluations on benchmark datasets at various data scales demonstrate that UnitY outperforms a single-pass speech-to-unit translation model by 2.5-4.2 ASR-BLEU with 2.83x decoding speed-up. We show that the proposed methods boost the performance even when predicting spectrogram in the second pass. However, predicting discrete units achieves 2.51x decoding speed-up compared to that case.

Keywords

Cite

@article{arxiv.2212.08055,
  title  = {UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units},
  author = {Hirofumi Inaguma and Sravya Popuri and Ilia Kulikov and Peng-Jen Chen and Changhan Wang and Yu-An Chung and Yun Tang and Ann Lee and Shinji Watanabe and Juan Pino},
  journal= {arXiv preprint arXiv:2212.08055},
  year   = {2023}
}

Comments

ACL 2023 (main conference)

R2 v1 2026-06-28T07:37:28.002Z