English

ComSL: A Composite Speech-Language Model for End-to-End Speech-to-Text Translation

Computation and Language 2023-10-17 v2 Sound Audio and Speech Processing

Abstract

Joint speech-language training is challenging due to the large demand for training data and GPU consumption, as well as the modality gap between speech and language. We present ComSL, a speech-language model built atop a composite architecture of public pretrained speech-only and language-only models and optimized data-efficiently for spoken language tasks. Particularly, we propose to incorporate cross-modality learning into transfer learning and conduct them simultaneously for downstream tasks in a multi-task learning manner. Our approach has demonstrated effectiveness in end-to-end speech-to-text translation tasks, achieving a new state-of-the-art average BLEU score of 31.5 on the multilingual speech to English text translation task for 21 languages, as measured on the public CoVoST2 evaluation set.

Keywords

Cite

@article{arxiv.2305.14838,
  title  = {ComSL: A Composite Speech-Language Model for End-to-End Speech-to-Text Translation},
  author = {Chenyang Le and Yao Qian and Long Zhou and Shujie Liu and Yanmin Qian and Michael Zeng and Xuedong Huang},
  journal= {arXiv preprint arXiv:2305.14838},
  year   = {2023}
}

Comments

NeurIPS 2023, Poster

R2 v1 2026-06-28T10:44:09.057Z