English

LLaST: Improved End-to-end Speech Translation System Leveraged by Large Language Models

Computation and Language 2024-07-23 v1

Abstract

We introduces LLaST, a framework for building high-performance Large Language model based Speech-to-text Translation systems. We address the limitations of end-to-end speech translation(E2E ST) models by exploring model architecture design and optimization techniques tailored for LLMs. Our approach includes LLM-based speech translation architecture design, ASR-augmented training, multilingual data augmentation, and dual-LoRA optimization. Our approach demonstrates superior performance on the CoVoST-2 benchmark and showcases exceptional scaling capabilities powered by LLMs. We believe this effective method will serve as a strong baseline for speech translation and provide insights for future improvements of the LLM-based speech translation framework. We release the data, code and models in https://github.com/openaudiolab/LLaST.

Keywords

Cite

@article{arxiv.2407.15415,
  title  = {LLaST: Improved End-to-end Speech Translation System Leveraged by Large Language Models},
  author = {Xi Chen and Songyang Zhang and Qibing Bai and Kai Chen and Satoshi Nakamura},
  journal= {arXiv preprint arXiv:2407.15415},
  year   = {2024}
}
R2 v1 2026-06-28T17:49:11.085Z