English

Using Songs to Improve Kazakh Automatic Speech Recognition

Audio and Speech Processing 2026-03-10 v3

Abstract

Developing automatic speech recognition (ASR) systems for low-resource languages is hindered by the scarcity of transcribed corpora. This proof-of-concept study explores songs as an unconventional yet promising data source for Kazakh ASR. We curate a dataset of 3,013 audio-text pairs (about 4.5 hours) from 195 songs by 36 artists, segmented at the lyric-line level. Using Whisper as the base recogniser, we fine-tune models under seven training scenarios involving Songs, Common Voice Corpus (CVC), and FLEURS, and evaluate them on three benchmarks: CVC, FLEURS, and Kazakh Speech Corpus 2 (KSC2). Results show that song-based fine-tuning improves performance over zero-shot baselines. For instance, Whisper Large-V3 Turbo trained on a mixture of Songs, CVC, and FLEURS achieves 27.6% normalised WER on CVC and 11.8% on FLEURS, while halving the error on KSC2 (39.3% vs. 81.2%) relative to the zero-shot model. Although these gains remain below those of models trained on the 1,100-hour KSC2 corpus, they demonstrate that even modest song-speech mixtures can yield meaningful adaptation improvements in low-resource ASR. The dataset is released on Hugging Face for research purposes under a gated, non-commercial licence.

Keywords

Cite

@article{arxiv.2603.00961,
  title  = {Using Songs to Improve Kazakh Automatic Speech Recognition},
  author = {Rustem Yeshpanov},
  journal= {arXiv preprint arXiv:2603.00961},
  year   = {2026}
}

Comments

10 pages, 7 tables, to appear in Proceedings of the 2026 Language Resources and Evaluation Conference

R2 v1 2026-07-01T10:57:45.120Z