English

The Multilingual TEDx Corpus for Speech Recognition and Translation

Computation and Language 2021-06-16 v2

Abstract

We present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a collection of audio recordings from TEDx talks in 8 source languages. We segment transcripts into sentences and align them to the source-language audio and target-language translations. The corpus is released along with open-sourced code enabling extension to new talks and languages as they become available. Our corpus creation methodology can be applied to more languages than previous work, and creates multi-way parallel evaluation sets. We provide baselines in multiple ASR and ST settings, including multilingual models to improve translation performance for low-resource language pairs.

Keywords

Cite

@article{arxiv.2102.01757,
  title  = {The Multilingual TEDx Corpus for Speech Recognition and Translation},
  author = {Elizabeth Salesky and Matthew Wiesner and Jacob Bremerman and Roldano Cattoni and Matteo Negri and Marco Turchi and Douglas W. Oard and Matt Post},
  journal= {arXiv preprint arXiv:2102.01757},
  year   = {2021}
}

Comments

Accepted to Interspeech 2021

R2 v1 2026-06-23T22:46:53.552Z