English

Updated Corpora and Benchmarks for Long-Form Speech Recognition

Computation and Language 2023-09-27 v1 Sound Audio and Speech Processing

Abstract

The vast majority of ASR research uses corpora in which both the training and test data have been pre-segmented into utterances. In most real-word ASR use-cases, however, test audio is not segmented, leading to a mismatch between inference-time conditions and models trained on segmented utterances. In this paper, we re-release three standard ASR corpora - TED-LIUM 3, Gigapeech, and VoxPopuli-en - with updated transcription and alignments to enable their use for long-form ASR research. We use these reconstituted corpora to study the train-test mismatch problem for transducers and attention-based encoder-decoders (AEDs), confirming that AEDs are more susceptible to this issue. Finally, we benchmark a simple long-form training for these models, showing its efficacy for model robustness under this domain shift.

Keywords

Cite

@article{arxiv.2309.15013,
  title  = {Updated Corpora and Benchmarks for Long-Form Speech Recognition},
  author = {Jennifer Drexler Fox and Desh Raj and Natalie Delworth and Quinn McNamara and Corey Miller and Migüel Jetté},
  journal= {arXiv preprint arXiv:2309.15013},
  year   = {2023}
}

Comments

Submitted to ICASSP 2024

R2 v1 2026-06-28T12:32:52.126Z