English

Self-Training for End-to-End Speech Translation

Computation and Language 2020-10-14 v2 Sound Audio and Speech Processing

Abstract

One of the main challenges for end-to-end speech translation is data scarcity. We leverage pseudo-labels generated from unlabeled audio by a cascade and an end-to-end speech translation model. This provides 8.3 and 5.7 BLEU gains over a strong semi-supervised baseline on the MuST-C English-French and English-German datasets, reaching state-of-the art performance. The effect of the quality of the pseudo-labels is investigated. Our approach is shown to be more effective than simply pre-training the encoder on the speech recognition task. Finally, we demonstrate the effectiveness of self-training by directly generating pseudo-labels with an end-to-end model instead of a cascade model.

Keywords

Cite

@article{arxiv.2006.02490,
  title  = {Self-Training for End-to-End Speech Translation},
  author = {Juan Pino and Qiantong Xu and Xutai Ma and Mohammad Javad Dousti and Yun Tang},
  journal= {arXiv preprint arXiv:2006.02490},
  year   = {2020}
}

Comments

INTERSPEECH 2020

R2 v1 2026-06-23T16:02:19.781Z