Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
Abstract
We introduce Balalaika, an open-source, data-centric pipeline for processing audio and producing prosody-aware annotations. It combines semantic VAD for context-preserving segmentation, multi-ASR ensembling with ROVER consensus decoding, while retaining optional word-level timestamps, followed by automatic quality and speaker-purity filtering. The text is further enriched with punctuation restoration, lexical stress and "\textipa{e}/\textipa{\H{e}}" normalization, and IPA phonemes. Using Balalaika, we build a 5.1k-hour multi-source Russian corpus with rich annotations, and show consistent gains under equalized training budgets for both speech denoising and TTS; ablations confirm complementary benefits of stress and punctuation and improved synthesis with stricter MOS filtering. The datasets are publicly available at \href{https://huggingface.co/collections/lab260/balalaika-dataset}{\underline{\textbf{HuggingFace}}}
Keywords
Cite
@article{arxiv.2507.13563,
title = {Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech},
author = {Kirill Borodin and Nikita Vasiliev and Vasiliy Kudryavtsev and Maxim Maslov and Mikhail Gorodnichev and Grach Mkrtchian},
journal= {arXiv preprint arXiv:2507.13563},
year = {2026}
}
Comments
The work is still in progress. Submitted to Interspeech 2026