English

Speech Wikimedia: A 77 Language Multilingual Speech Dataset

Artificial Intelligence 2023-08-31 v1 Machine Learning

Abstract

The Speech Wikimedia Dataset is a publicly available compilation of audio with transcriptions extracted from Wikimedia Commons. It includes 1780 hours (195 GB) of CC-BY-SA licensed transcribed speech from a diverse set of scenarios and speakers, in 77 different languages. Each audio file has one or more transcriptions in different languages, making this dataset suitable for training speech recognition, speech translation, and machine translation models.

Keywords

Cite

@article{arxiv.2308.15710,
  title  = {Speech Wikimedia: A 77 Language Multilingual Speech Dataset},
  author = {Rafael Mosquera Gómez and Julián Eusse and Juan Ciro and Daniel Galvez and Ryan Hileman and Kurt Bollacker and David Kanter},
  journal= {arXiv preprint arXiv:2308.15710},
  year   = {2023}
}

Comments

Data-Centric Machine Learning Workshop at the International Machine Learning Conference 2023 (ICML)

R2 v1 2026-06-28T12:07:57.557Z