English

Recent Advances in Direct Speech-to-text Translation

Computation and Language 2023-06-21 v1 Audio and Speech Processing

Abstract

Recently, speech-to-text translation has attracted more and more attention and many studies have emerged rapidly. In this paper, we present a comprehensive survey on direct speech translation aiming to summarize the current state-of-the-art techniques. First, we categorize the existing research work into three directions based on the main challenges -- modeling burden, data scarcity, and application issues. To tackle the problem of modeling burden, two main structures have been proposed, encoder-decoder framework (Transformer and the variants) and multitask frameworks. For the challenge of data scarcity, recent work resorts to many sophisticated techniques, such as data augmentation, pre-training, knowledge distillation, and multilingual modeling. We analyze and summarize the application issues, which include real-time, segmentation, named entity, gender bias, and code-switching. Finally, we discuss some promising directions for future work.

Keywords

Cite

@article{arxiv.2306.11646,
  title  = {Recent Advances in Direct Speech-to-text Translation},
  author = {Chen Xu and Rong Ye and Qianqian Dong and Chengqi Zhao and Tom Ko and Mingxuan Wang and Tong Xiao and Jingbo Zhu},
  journal= {arXiv preprint arXiv:2306.11646},
  year   = {2023}
}

Comments

An expanded version of the paper accepted by IJCAI2023 survey track

R2 v1 2026-06-28T11:09:49.423Z