English

ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets

Sound 2024-06-14 v1 Computation and Language Audio and Speech Processing

Abstract

ML-SUPERB evaluates self-supervised learning (SSL) models on the tasks of language identification and automatic speech recognition (ASR). This benchmark treats the models as feature extractors and uses a single shallow downstream model, which can be fine-tuned for a downstream task. However, real-world use cases may require different configurations. This paper presents ML-SUPERB~2.0, which is a new benchmark for evaluating pre-trained SSL and supervised speech models across downstream models, fine-tuning setups, and efficient model adaptation approaches. We find performance improvements over the setup of ML-SUPERB. However, performance depends on the downstream model design. Also, we find large performance differences between languages and datasets, suggesting the need for more targeted approaches to improve multilingual ASR performance.

Keywords

Cite

@article{arxiv.2406.08641,
  title  = {ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets},
  author = {Jiatong Shi and Shih-Heng Wang and William Chen and Martijn Bartelds and Vanya Bannihatti Kumar and Jinchuan Tian and Xuankai Chang and Dan Jurafsky and Karen Livescu and Hung-yi Lee and Shinji Watanabe},
  journal= {arXiv preprint arXiv:2406.08641},
  year   = {2024}
}

Comments

Accepted by Interspeech 2024

R2 v1 2026-06-28T17:03:47.594Z