English

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data

Computation and Language 2025-06-03 v1 Audio and Speech Processing

Abstract

This paper reports on the development of a large-scale speech recognition model, Whale. Similar to models such as Whisper and OWSM, Whale leverages both a large model size and a diverse, extensive dataset. Whale's architecture integrates w2v-BERT self-supervised model, an encoder-decoder backbone built on E-Branchformer, and a joint CTC-attention decoding strategy. The training corpus comprises varied speech data, of not only public corpora but also in-house data, thereby enhancing the model's robustness to different speaking styles and acoustic conditions. Through evaluations on multiple benchmarks, Whale achieved comparable performance to existing models. In particular, it achieves a word error rate of 2.4% on the Librispeech test-clean set and a character error rate of 3.4% on the CSJ eval3 set, outperforming Whisper large-v3 and OWSM v3.1.

Keywords

Cite

@article{arxiv.2506.01439,
  title  = {Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data},
  author = {Yosuke Kashiwagi and Hayato Futami and Emiru Tsunoo and Satoshi Asakawa},
  journal= {arXiv preprint arXiv:2506.01439},
  year   = {2025}
}
R2 v1 2026-07-01T02:53:58.245Z