English

Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech

Computation and Language 2026-05-05 v2 Sound Audio and Speech Processing

Abstract

We introduce Balalaika, an open-source, data-centric pipeline for processing audio and producing prosody-aware annotations. It combines semantic VAD for context-preserving segmentation, multi-ASR ensembling with ROVER consensus decoding, while retaining optional word-level timestamps, followed by automatic quality and speaker-purity filtering. The text is further enriched with punctuation restoration, lexical stress and "\textipa{e}/\textipa{\H{e}}" normalization, and IPA phonemes. Using Balalaika, we build a 5.1k-hour multi-source Russian corpus with rich annotations, and show consistent gains under equalized training budgets for both speech denoising and TTS; ablations confirm complementary benefits of stress and punctuation and improved synthesis with stricter MOS filtering. The datasets are publicly available at \href{https://huggingface.co/collections/lab260/balalaika-dataset}{\underline{\textbf{HuggingFace}}}

Keywords

Cite

@article{arxiv.2507.13563,
  title  = {Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech},
  author = {Kirill Borodin and Nikita Vasiliev and Vasiliy Kudryavtsev and Maxim Maslov and Mikhail Gorodnichev and Grach Mkrtchian},
  journal= {arXiv preprint arXiv:2507.13563},
  year   = {2026}
}

Comments

The work is still in progress. Submitted to Interspeech 2026

R2 v1 2026-07-01T04:07:04.753Z