English
Related papers

Related papers: Balalaika: Data-Centric, Prosody-Aware Annotation …

200 papers

Vision-Language-Action (VLA) systems have shown strong potential for language-driven robotic manipulation. However, scaling them to long-horizon tasks remains challenging. Existing pipelines typically separate data collection, policy…

With the emergence of audio-language models, constructing large-scale paired audio-language datasets has become essential yet challenging for model development, primarily due to the time-intensive and labour-heavy demands involved. While…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-02 Jisheng Bai , Haohe Liu , Mou Wang , Dongyuan Shi , Wenwu Wang , Mark D. Plumbley , Woon-Seng Gan , Jianfeng Chen

Prosodic boundary plays an important role in text-to-speech synthesis (TTS) in terms of naturalness and readability. However, the acquisition of prosodic boundary labels relies on manual annotation, which is costly and time-consuming. In…

Sound · Computer Science 2022-06-17 Ziqian Dai , Jianwei Yu , Yan Wang , Nuo Chen , Yanyao Bian , Guangzhi Li , Deng Cai , Dong Yu

We propose several improvements to the speech recognition evaluation. First, we propose a string alignment algorithm that supports both multi-reference labeling, arbitrary-length insertions and better word alignment. This is especially…

Computation and Language · Computer Science 2026-01-30 Oleg Sedukhin , Andrey Kostin

The Tajik language, written in Cyrillic script, remains severely under-resourced in terms of publicly available natural language processing (NLP) toolkits, hindering both linguistic research and applied development. This paper introduces…

Computation and Language · Computer Science 2026-05-29 Mullosharaf K. Arabov

Medical audio data is difficult to collect due to privacy regulations and high annotation costs arising from domain expertise. Thus, existing benchmarks tend to underrepresent complex medical audio scenarios. To address this challenge, we…

The HuggingFace Datasets Hub hosts thousands of datasets, offering exciting opportunities for language model training and evaluation. However, datasets for a specific task type often have different schemas, making harmonization challenging.…

Computation and Language · Computer Science 2023-05-17 Damien Sileo

Despite advances in language and speech technologies, no open-source system enables full speech-to-speech, multi-turn dialogue with integrated tool use and agentic reasoning. We introduce AURA (Agent for Understanding, Reasoning, and…

Artificial Intelligence · Computer Science 2025-07-01 Leander Melroy Maben , Gayathri Ganesh Lakshmy , Srijith Radhakrishnan , Siddhant Arora , Shinji Watanabe

Project Euphonia, a Google initiative, is dedicated to improving automatic speech recognition (ASR) of disordered speech. A central objective of the project is to create a large, high-quality, and diverse speech corpus. This report…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-17 Pan-Pan Jiang , Jimmy Tobin , Katrin Tomanek , Robert L. MacDonald , Katie Seaver , Richard Cave , Marilyn Ladewig , Rus Heywood , Jordan R. Green

Automatic Speech Recognition (ASR) for adults' speeches has made significant progress by employing deep neural network (DNN) models recently, but improvement in children's speech is still unsatisfactory due to children's speech's distinct…

Computation and Language · Computer Science 2024-06-27 Dancheng Liu , Jinjun Xiong

In this paper, we introduce Kathaka, a model trained with a novel two-stage training process for neural speech synthesis with contextually appropriate prosody. In Stage I, we learn a prosodic distribution at the sentence level from…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-05 Sri Karlapati , Ammar Abbas , Zack Hodari , Alexis Moinet , Arnaud Joly , Penny Karanasou , Thomas Drugman

Automatic fluency assessment (AFA) remains challenging, particularly in capturing speech rhythm, pauses, and disfluencies in non-native speakers. We introduce a chunk-based approach integrating self-supervised learning (SSL) models…

Computation and Language · Computer Science 2025-06-27 Papa Séga Wade , Mihai Andries , Ioannis Kanellos , Thierry Moudenc

Vision-language pre-training (VLP) offers unique advantages for surgery by aligning language with surgical videos, enabling workflow understanding and transfer across tasks without relying on expert-labeled datasets. However, progress in…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Alejandra Perez , Chinedu Nwoye , Ramtin Raji Kermani , Omid Mohareri , Muhammad Abdullah Jamal

Modern machine learning systems rely on complex data engineering workflows to extract, transform, and load (ELT) data into production pipelines. However, constructing these pipelines remains time-consuming and requires substantial expertise…

Software Engineering · Computer Science 2026-03-24 Rohan Siva , Kai Cheung , Lichi Li , Ganesh Sundaram

Unlike the Open Domain Question Answering (ODQA) setting, the conversational (ODConvQA) domain has received limited attention when it comes to reevaluating baselines for both efficiency and effectiveness. In this paper, we study the…

Computation and Language · Computer Science 2023-10-24 Andrei C. Coman , Gianni Barlacchi , Adrià de Gispert

This paper presents a novel Dialectal Sound and Vowelization Recovery framework, designed to recognize borrowed and dialectal sounds within phonologically diverse and dialect-rich languages, that extends beyond its standard orthographic…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-06 Yassine El Kheir , Hamdy Mubarak , Ahmed Ali , Shammur Absar Chowdhury

The large-scale digitization of historical archives has created a paradox: "dark data"-digital objects lacking metadata for retrieval. Manual archival description is slow and expensive, limiting discovery and reuse. We propose Vidya, a…

Digital Libraries · Computer Science 2026-05-19 Cloter Migliorini Filho , Julia Graciela Machado , Edson Armando Silva , Marcella Scoczynski

In this paper, we present a novel series of Russian information retrieval datasets constructed from the "Did you know..." section of Russian Wikipedia. Our datasets support a range of retrieval tasks, including fact-checking,…

Information Retrieval · Computer Science 2025-11-10 Grigory Kovalev , Natalia Loukachevitch , Mikhail Tikhomirov , Olga Babina , Pavel Mamaev

Automatic speech recognition (ASR) systems often degrade on accented speech because acoustic-phonetic and prosodic shifts induce a mismatch to training data, making labeled accent adaptation costly. However, common pseudo-label selection…

Computation and Language · Computer Science 2026-02-17 Ligong Lei , Wenwen Lu , Xudong Pang , Zaokere Kadeer , Aishan Wumaier

This paper introduces a human-in-the-loop (HITL) data annotation pipeline to generate high-quality, large-scale speech datasets. The pipeline combines human and machine advantages to more quickly, accurately, and cost-effectively annotate…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-06 Mingkuan Liu , Chi Zhang , Hua Xing , Chao Feng , Monchu Chen , Judith Bishop , Grace Ngapo