中文

WorldSpeech:全球多语言语音语料库

计算与语言 2026-05-12 v1 人工智能 机器学习

摘要

自动语音识别(ASR)在高资源语言(即拥有大量配对音频-文字数据的语言)上表现良好,但由于大多数语言缺乏公开可获取的对齐数据,其准确率会急剧下降。为此,我们介绍 WorldSpeech,一个由 24 kHz 组成的多语言语音语料库,包含 65000 小时的对齐音频-文字数据,覆盖 76 种语言,数据来源于议会记录、国际广播及公共领域有声书等多样化的公开来源。对于 37 种语言,WorldSpeech 提供超过 200 小时的对齐语音,其中 28 种超过 500 小时,24 种超过 1000 小时。在 WorldSpeech 上微调现有 ASR 模型后,平均相对词错误率(WER)在 11 种类型多样的语言上降低了 63.5%。

关键词

引用

@article{arxiv.2605.09167,
  title  = {WorldSpeech: A Multilingual Speech Corpus from Around the World},
  author = {Antonis Asonitis and Luca A. Lanzendörfer and Frédéric Berdoz and Roger Wattenhofer},
  journal= {arXiv preprint arXiv:2605.09167},
  year   = {2026}
}