English

Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation

Computation and Language 2024-06-28 v2 Sound Audio and Speech Processing

Abstract

This paper describes our system submission to the International Conference on Spoken Language Translation (IWSLT 2024) for Irish-to-English speech translation. We built end-to-end systems based on Whisper, and employed a number of data augmentation techniques, such as speech back-translation and noise augmentation. We investigate the effect of using synthetic audio data and discuss several methods for enriching signal diversity.

Keywords

Cite

@article{arxiv.2406.17363,
  title  = {Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation},
  author = {Yasmin Moslem},
  journal= {arXiv preprint arXiv:2406.17363},
  year   = {2024}
}

Comments

IWSLT 2024