English
Related papers

Related papers: Balalaika: Data-Centric, Prosody-Aware Annotation …

200 papers

This paper presents Praaline, an open-source software system for managing, annotating, analysing and visualising speech corpora. Researchers working with speech corpora are often faced with multiple tools and formats, and they need to work…

Computation and Language · Computer Science 2018-02-09 George Christodoulides

RuASD (Russian AntiSpoofing Dataset) is a dedicated, reproducible benchmark for Russian-language speech anti-spoofing designed to evaluate both in-domain discrimination and robustness to deployment-style distribution shifts. It combines a…

Sound · Computer Science 2026-04-28 Ksenia Lysikova , Kirill Borodin , Grach Mkrtchian

We present RUSLAN -- a new open Russian spoken language corpus for the text-to-speech task. RUSLAN contains 22200 audio samples with text annotations -- more than 31 hours of high-quality speech of one person -- being the largest annotated…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-28 Lenar Gabdrakhmanov , Rustem Garaev , Evgenii Razinkov

This paper introduces a novel Russian speech dataset called Golos, a large corpus suitable for speech research. The dataset mainly consists of recorded audio files manually annotated on the crowd-sourcing platform. The total duration of the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-21 Nikolay Karpov , Alexander Denisenko , Fedor Minkin

We present CLASSLA-Stanza, a pipeline for automatic linguistic annotation of the South Slavic languages, which is based on the Stanza natural language processing pipeline. We describe the main improvements in CLASSLA-Stanza with respect to…

Computation and Language · Computer Science 2023-08-14 Luka Terčon , Nikola Ljubešić

Recent progress in speech processing has highlighted that high-quality performance across languages requires substantial training data for each individual language. While existing multilingual datasets cover many languages, they often…

Computation and Language · Computer Science 2025-10-28 Samuel Pfisterer , Florian Grötschla , Luca A. Lanzendörfer , Florian Yan , Roger Wattenhofer

We introduce T-pro 2.0, an open-weight Russian LLM for hybrid reasoning and efficient inference. The model supports direct answering and reasoning-trace generation, using a Cyrillic-dense tokenizer and an adapted EAGLE speculative-decoding…

In this work, we showcase a cost-effective method for generating training data for speech processing tasks. First, we transcribe unlabeled speech using a state-of-the-art Automatic Speech Recognition (ASR) model. Next, we align generated…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-19 Taras Sereda

Defining psycholinguistic characteristics in written texts is a task gaining increasing attention from researchers. One of the most widely used tools in the current field is Linguistic Inquiry and Word Count (LIWC) that originally was…

Computation and Language · Computer Science 2026-01-29 Elina Sigdel , Anastasia Panfilova

Many archival recordings of speech from endangered languages remain unannotated and inaccessible to community members and language learning programs. One bottleneck is the time-intensive nature of annotation. An even narrower bottleneck…

Computation and Language · Computer Science 2022-04-26 Nay San , Martijn Bartelds , Tolúlopé Ògúnrèmí , Alison Mount , Ruben Thompson , Michael Higgins , Roy Barker , Jane Simpson , Dan Jurafsky

Paralinguistic vocalizations-including non-verbal sounds like laughter and breathing, as well as lexicalized interjections such as "uhm" and "oh"-are integral to natural spoken communication. Despite their importance in conveying affect,…

Sound · Computer Science 2025-08-07 Huan Liao , Qinke Ni , Yuancheng Wang , Yiheng Lu , Haoyue Zhan , Pengyuan Xie , Qiang Zhang , Zhizheng Wu

This paper presents an overview of rule-based system for automatic accentuation and phonemic transcription of Russian texts for speech connected tasks, such as Automatic Speech Recognition (ASR). Two parts of the developed system,…

Computation and Language · Computer Science 2024-10-07 Olga Iakovenko , Ivan Bondarenko , Mariya Borovikova , Daniil Vodolazsky

The adaptation of Large-Scale Language Models (LLMs) to specific domains depends on high-quality fine-tuning datasets, particularly in instructional format (e.g., Question-Answer - Q&A). However, generating these datasets, particularly from…

Machine Learning · Computer Science 2026-01-22 Alex Echeverria , Sávio Salvarino Teles de Oliveira , Fernando Marques Federson

We present a new data set for speech emotion recognition (SER) tasks called Dusha. The corpus contains approximately 350 hours of data, more than 300 000 audio recordings with Russian speech and their transcripts. Therefore it is the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-26 Vladimir Kondratenko , Artem Sokolov , Nikolay Karpov , Oleg Kutuzov , Nikita Savushkin , Fyodor Minkin

Current anti-spoofing systems remain vulnerable to expressive and emotional synthetic speech, since they rarely leverage prosody as a discriminative cue. Prosody is central to human expressiveness and emotion, and humans instinctively use…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-29 Aurosweta Mahapatra , Ismail Rasim Ulgen , Berrak Sisman

In expressive and controllable Text-to-Speech (TTS), explicit prosodic features significantly improve the naturalness and controllability of synthesised speech. However, manual prosody annotation is labor-intensive and inconsistent. To…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Jinzuomu Zhong , Yang Li , Hui Huang , Korin Richmond , Jie Liu , Zhiba Su , Jing Guo , Benlai Tang , Fengjie Zhu

Data scarcity and noise are important issues in industrial applications of machine learning. However, it is often challenging to devise a scalable and generalized approach to address the fundamental distributional and semantic properties of…

Machine Learning · Computer Science 2021-12-08 Youngjune Lee , Oh Joon Kwon , Haeju Lee , Joonyoung Kim , Kangwook Lee , Kee-Eung Kim

ParlaSpeech is a collection of spoken parliamentary corpora currently spanning four Slavic languages - Croatian, Czech, Polish and Serbian - all together 6 thousand hours in size. The corpora were built in an automatic fashion from the…

Computation and Language · Computer Science 2026-04-16 Nikola Ljubešić , Peter Rupnik , Ivan Porupski , Taja Kuzman Pungeršek

We propose to address online speaker diarization as a combination of incremental clustering and local diarization applied to a rolling buffer updated every 500ms. Every single step of the proposed pipeline is designed to take full advantage…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-15 Juan M. Coria , Hervé Bredin , Sahar Ghannay , Sophie Rosset

Voice conversion (VC) research traditionally depends on scripted or acted speech, which lacks the natural spontaneity of real-life conversations. While natural speech data is limited for VC, our study focuses on filling in this gap. We…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-10 Ali N. Salman , Zongyang Du , Shreeram Suresh Chandra , Ismail Rasim Ulgen , Carlos Busso , Berrak Sisman
‹ Prev 1 2 3 10 Next ›