English
Related papers

Related papers: WhispEar: A Bi-directional Framework for Scaling W…

200 papers

Speech enhancement is critical for improving speech intelligibility and quality in various audio devices. In recent years, deep learning-based methods have significantly improved speech enhancement performance, but they often come with a…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-08 Xiang Hao , Chenxiang Ma , Qu Yang , Jibin Wu , Kay Chen Tan

Automatic Speech Recognition (ASR) systems have progressed significantly in their performance on adult speech data; however, transcribing child speech remains challenging due to the acoustic differences in the characteristics of child and…

Computation and Language · Computer Science 2023-11-10 Andrei Barcovschi , Rishabh Jain , Peter Corcoran

This paper proposes a new architecture for speaker adaptation of multi-speaker neural-network speech synthesis systems, in which an unseen speaker's voice can be built using a relatively small amount of speech data without transcriptions.…

Audio and Speech Processing · Electrical Eng. & Systems 2018-08-21 Hieu-Thi Luong , Junichi Yamagishi

The Open Whisper-style Speech Model (OWSM) series was introduced to achieve full transparency in building advanced speech-to-text (S2T) foundation models. To this end, OWSM models are trained on 25 public speech datasets, which are…

Computation and Language · Computer Science 2024-06-14 Jinchuan Tian , Yifan Peng , William Chen , Kwanghee Choi , Karen Livescu , Shinji Watanabe

A universal audio representation should capture fine-grained speech cues and high-level semantics for environmental sounds and music in a single encoder. Existing encoders often excel in one domain but degrade in others. We propose…

Sound · Computer Science 2026-03-10 Yuxuan Chen , Peize He , Haoyuan Yu , Junzi Zhang

Sparse Autoencoders (SAEs) are powerful tools for interpreting neural representations, yet their use in audio remains underexplored. We train SAEs across all encoder layers of Whisper and HuBERT, provide an extensive evaluation of their…

This paper presents Whisper, a fast and reliable protocol to flood small amounts of data into a multi-hop network. Whisper makes use of synchronous transmissions, a technique first introduced by the Glossy flooding protocol. In contrast to…

Networking and Internet Architecture · Computer Science 2021-05-28 Martina Brachmann , Olaf Landsiedel , Diana Göhringer , Silvia Santini

Generative Universal Speech Enhancement (USE) methods aim to leverage generative models to improve speech quality under various types of distortions. However, existing generative speech enhancement methods often suffer from semantic…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-07 Xingchen Li , Hanke Xie , Ziqian Wang , Zihan Zhang , Longshuai Xiao , Shuai Wang , Lei Xie

Conditional sound separation in multi-source audio mixtures without having access to single source sound data during training is a long standing challenge. Existing mix-and-separate based methods suffer from significant performance drop…

Sound · Computer Science 2024-04-03 Tanvir Mahmud , Saeed Amizadeh , Kazuhito Koishida , Diana Marculescu

Recent years have witnessed significant progress in multilingual automatic speech recognition (ASR), driven by the emergence of end-to-end (E2E) models and the scaling of multilingual datasets. Despite that, two main challenges persist in…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Zheshu Song , Jianheng Zhuo , Yifan Yang , Ziyang Ma , Shixiong Zhang , Xie Chen

We introduce a new cross-modal fusion technique designed for generative error correction in automatic speech recognition (ASR). Our methodology leverages both acoustic information and external linguistic representations to generate accurate…

Automatic Speech Recognition (ASR) systems are known to exhibit difficulties when transcribing children's speech. This can mainly be attributed to the absence of large children's speech corpora to train robust ASR models and the resulting…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-22 Jenthe Thienpondt , Kris Demuynck

Reliable transcription of child-adult conversations in clinical settings is crucial for diagnosing developmental disorders like Autism. Recent advances in deep learning and availability of large scale transcribed data has led to development…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-15 Aditya Ashvin , Rimita Lahiri , Aditya Kommineni , Somer Bishop , Catherine Lord , Sudarsana Reddy Kadiri , Shrikanth Narayanan

This paper examines the integration of real-time talking-head generation for interviewer training, focusing on overcoming challenges in Audio Feature Extraction (AFE), which often introduces latency and limits responsiveness in real-time…

Whisper, despite being trained on 680K hours of web-scaled audio data, faces difficulty in recognising rare words like domain-specific terms, with a solution being contextual biasing through prompting. To improve upon this method, in this…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-19 Yash Jogi , Vaibhav Aggarwal , Shabari S Nair , Yash Verma , Aayush Kubba

Generative models have excelled in audio tasks using approaches such as language models, diffusion, and flow matching. However, existing generative approaches for speech enhancement (SE) face notable challenges: language model-based methods…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-28 Ziqian Wang , Zikai Liu , Xinfa Zhu , Yike Zhu , Mingshuai Liu , Jun Chen , Longshuai Xiao , Chao Weng , Lei Xie

Recently, we made available WeNet, a production-oriented end-to-end speech recognition toolkit, which introduces a unified two-pass (U2) framework and a built-in runtime to address the streaming and non-streaming decoding modes in a single…

Sound · Computer Science 2022-07-06 Binbin Zhang , Di Wu , Zhendong Peng , Xingchen Song , Zhuoyuan Yao , Hang Lv , Lei Xie , Chao Yang , Fuping Pan , Jianwei Niu

Direct speech-to-speech translation (S2ST) with discrete self-supervised representations has achieved remarkable accuracy, but is unable to preserve the speaker timbre of the source speech. Meanwhile, the scarcity of high-quality…

Sound · Computer Science 2024-07-22 Yongqi Wang , Jionghao Bai , Rongjie Huang , Ruiqi Li , Zhiqing Hong , Zhou Zhao

End-to-end speech-to-speech translation (S2ST) systems typically struggle with a critical data bottleneck: the scarcity of parallel speech-to-speech corpora. To overcome this, we introduce RosettaSpeech, a novel zero-shot framework trained…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-17 Zhisheng Zheng , Xiaohang Sun , Tuan Dinh , Abhishek Yanamandra , Abhinav Jain , Zhu Liu , Sunil Hadap , Vimal Bhat , Manoj Aggarwal , Gerard Medioni , David Harwath

Recent speech-to-speech (S2S) models generate intelligible speech but still lack natural expressiveness, largely due to the absence of a reliable evaluation metric. Existing approaches, such as subjective MOS ratings, low-level acoustic…

Sound · Computer Science 2025-10-24 Zhiyu Lin , Jingwen Yang , Jiale Zhao , Meng Liu , Sunzhu Li , Benyou Wang