English

End-to-End Spoken Language Translation

Computation and Language 2019-04-25 v1 Sound Audio and Speech Processing

Abstract

In this paper, we address the task of spoken language understanding. We present a method for translating spoken sentences from one language into spoken sentences in another language. Given spectrogram-spectrogram pairs, our model can be trained completely from scratch to translate unseen sentences. Our method consists of a pyramidal-bidirectional recurrent network combined with a convolutional network to output sentence-level spectrograms in the target language. Empirically, our model achieves competitive performance with state-of-the-art methods on multiple languages and can generalize to unseen speakers.

Keywords

Cite

@article{arxiv.1904.10760,
  title  = {End-to-End Spoken Language Translation},
  author = {Michelle Guo and Albert Haque and Prateek Verma},
  journal= {arXiv preprint arXiv:1904.10760},
  year   = {2019}
}

Comments

Technical Report. Stanford University, 2017. arXiv admin note: text overlap with arXiv:1804.00047

R2 v1 2026-06-23T08:48:13.907Z