English
Related papers

Related papers: Balalaika: Data-Centric, Prosody-Aware Annotation …

200 papers

This paper introduces a high-quality rich annotated Mandarin conversational (RAMC) speech dataset called MagicData-RAMC. The MagicData-RAMC corpus contains 180 hours of conversational speech data recorded from native speakers of Mandarin…

Computation and Language · Computer Science 2022-04-01 Zehui Yang , Yifan Chen , Lei Luo , Runyan Yang , Lingxuan Ye , Gaofeng Cheng , Ji Xu , Yaohui Jin , Qingqing Zhang , Pengyuan Zhang , Lei Xie , Yonghong Yan

Code-switching, the alternation between two or more languages within communication, poses great challenges for Automatic Speech Recognition (ASR) systems. Existing models and datasets are limited in their ability to effectively handle these…

Sound · Computer Science 2025-11-14 Yupei Li , Zifan Wei , Heng Yu , Jiahao Xue , Huichi Zhou , Björn W. Schuller

The labelling of speech corpora is a laborious and time-consuming process. The ProsoBeast Annotation Tool seeks to ease and accelerate this process by providing an interactive 2D representation of the prosodic landscape of the data, in…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-16 Branislav Gerazov , Michael Wagner

A voice AI agent that blends seamlessly into daily life would interact with humans in an autonomous, real-time, and emotionally expressive manner. Rather than merely reacting to commands, it would continuously listen, reason, and respond…

Artificial Intelligence · Computer Science 2025-05-06 Yemin Shi , Yu Shu , Siwei Dong , Guangyi Liu , Jaward Sesay , Jingwen Li , Zhiting Hu

The advancement of automatic speech recognition (ASR) has been largely enhanced by extensive datasets in high-resource languages, while languages such as Hungarian remain underrepresented due to limited spontaneous and conversational…

Computation and Language · Computer Science 2026-01-15 Máté Gedeon , Piroska Zsófia Barta , Péter Mihajlik , Tekla Etelka Gráczi , Anna Kohári , Katalin Mády

Toxic speech detection has become a crucial challenge in maintaining safe online communication environments. However, existing approaches to toxic speech detection often neglect the contribution of paralinguistic cues, such as emotion,…

Sound · Computer Science 2026-05-18 Zhongjie Ba , Liang Yi , Peng Cheng , Qingcao Li , Qinglong Wang , Li Lu

Acquiring large-scale emotional speech data with strong consistency remains a challenge for speech synthesis. This paper presents MIKU-PAL, a fully automated multimodal pipeline for extracting high-consistency emotional speech from…

Sound · Computer Science 2025-10-01 Yifan Cheng , Ruoyi Zhang , Jiatong Shi

Despite significant advances in speech processing, Portuguese remains under-resourced due to the scarcity of public, large-scale, and high-quality datasets. To address this gap, we present a new dataset, named TAGARELA, composed of over…

Clinical communication is central to patient outcomes, yet large-scale human annotation of patient-provider conversation remains labor-intensive, inconsistent, and difficult to scale. Existing approaches based on large language models…

In this work, we introduce a paralinguistic supervision paradigm for low-resource multilingual speech emotion recognition (LRM-SER) that leverages non-verbal vocalizations to exploit prosody-centric emotion cues. Unlike conventional SER…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-24 Girish , Mohd Mujtaba Akhtar , Muskaan Singh

Neural vocoders are central to speech synthesis; despite their success, most still suffer from limited prosody modeling and inaccurate phase reconstruction. We propose a vocoder that introduces prosody-guided harmonic attention to enhance…

Sound · Computer Science 2026-01-22 Mohammed Salah Al-Radhi , Riad Larbi , Mátyás Bartalis , Géza Németh

The lack of labeled second language (L2) speech data is a major challenge in designing mispronunciation detection models. We introduce SpeechBlender - a fine-grained data augmentation pipeline for generating mispronunciation errors to…

Sound · Computer Science 2023-07-13 Yassine El Kheir , Shammur Absar Chowdhury , Ahmed Ali , Hamdy Mubarak , Shazia Afzal

This work introduces the first framework for reconstructing surgical dialogue from unstructured real-world recordings, which is crucial for characterizing teaching tasks. In surgical training, the formative verbal feedback that trainers…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-03 Firdavs Nasriddinov , Rafal Kocielnik , Arushi Gupta , Cherine Yang , Elyssa Wong , Anima Anandkumar , Andrew Hung

Speech synthesis is crucial for human-computer interaction, enabling natural and intuitive communication. However, existing datasets involve high construction costs due to manual annotation and suffer from limited character diversity,…

Computation and Language · Computer Science 2025-04-22 Xiang Li , Duyi Pan , Hongru Xiao , Jiale Han , Jing Tang , Jiabao Ma , Wei Wang , Bo Cheng

Radiology reports contain rich clinical information that can be used to train imaging models without relying on costly manual annotation. However, existing approaches face critical limitations: rule-based methods struggle with linguistic…

In this study, we introduce YODAS (YouTube-Oriented Dataset for Audio and Speech), a large-scale, multilingual dataset comprising currently over 500k hours of speech data in more than 100 languages, sourced from both labeled and unlabeled…

Computation and Language · Computer Science 2024-06-04 Xinjian Li , Shinnosuke Takamichi , Takaaki Saeki , William Chen , Sayaka Shiota , Shinji Watanabe

Speech synthesis has recently seen significant improvements in fidelity, driven by the advent of neural vocoders and neural prosody generators. However, these systems lack intuitive user controls over prosody, making them unable to rectify…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-13 Max Morrison , Zeyu Jin , Justin Salamon , Nicholas J. Bryan , Gautham J. Mysore

Modern Automatic Speech Recognition (ASR) systems can achieve high performance in terms of recognition accuracy. However, a perfectly accurate transcript still can be challenging to read due to disfluency, filter words, and other errata…

Computation and Language · Computer Science 2021-02-23 Junwei Liao , Yu Shi , Ming Gong , Linjun Shou , Sefik Eskimez , Liyang Lu , Hong Qu , Michael Zeng

Audiovisual active speaker detection (ASD) in egocentric recordings is challenged by frequent occlusions, motion blur, and audio interference, which undermine the discernability of temporal synchrony between lip movement and speech.…

Multimedia · Computer Science 2025-08-15 Jason Clarke , Yoshihiko Gotoh , Stefan Goetze

Attention-based models have been widely used in many areas, such as computer vision and natural language processing. However, relevant applications in time series classification (TSC) have not been explored deeply yet, causing a significant…

Machine Learning · Computer Science 2022-07-18 Bowen Zhao , Huanlai Xing , Xinhan Wang , Fuhong Song , Zhiwen Xiao
‹ Prev 1 3 4 5 6 7 10 Next ›