English
Related papers

Related papers: Speech Separation with Pretrained Frontend to Mini…

200 papers

Because the performance of speech separation is excellent for speech in which two speakers completely overlap, research attention has been shifted to dealing with more realistic scenarios. However, domain mismatch between training/test…

Sound · Computer Science 2022-06-22 Fan-Lin Wang , Hung-Shin Lee , Yu Tsao , Hsin-Min Wang

Learning on a massive amount of speech corpus leads to the recent success of many self-supervised speech models. With knowledge distillation, these models may also benefit from the knowledge encoded by language models that are pre-trained…

Computation and Language · Computer Science 2023-03-08 Jinjie Ni , Yukun Ma , Wen Wang , Qian Chen , Dianwen Ng , Han Lei , Trung Hieu Nguyen , Chong Zhang , Bin Ma , Erik Cambria

Recent breakthroughs in multi-talker ASR (MT-ASR) and speaker diarization (SD) rely on synthetic data to mitigate the scarcity of large-scale conversational recordings, yet the impact of specific simulation choices remains poorly…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-18 Alexander Polok , Ivan Medennikov , Jan Černocký , Shinji Watanabe , Lukáš Burget , Samuele Cornell

Automatic speech recognition (ASR) in multimedia content is one of the promising applications, but speech data in this kind of content are frequently mixed with background music, which is harmful for the performance of ASR. In this study,…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-28 Jeongwoo Woo , Masato Mimura , Kazuyoshi Yoshii , Tatsuya Kawahara

Language models achieve impressive performance on a variety of knowledge, language, and reasoning tasks due to the scale and diversity of pretraining data available. The standard training recipe is a two-stage paradigm: pretraining first on…

Computation and Language · Computer Science 2026-03-20 Skyler Seto , Pierre Ablin , Anastasiia Filippova , Jiayuan Ye , Louis Bethune , Angelos Katharopoulos , David Grangier

Automatic Speech Understanding (ASU) leverages the power of deep learning models for accurate interpretation of human speech, leading to a wide range of speech applications that enrich the human experience. However, training a robust ASU…

Sound · Computer Science 2023-06-14 Tiantian Feng , Digbalay Bose , Xuan Shi , Shrikanth Narayanan

In a multi-channel separation task with multiple speakers, we aim to recover all individual speech signals from the mixture. In contrast to single-channel approaches, which rely on the different spectro-temporal characteristics of the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-11 Kristina Tesch , Timo Gerkmann

Conditional sound separation in multi-source audio mixtures without having access to single source sound data during training is a long standing challenge. Existing mix-and-separate based methods suffer from significant performance drop…

Sound · Computer Science 2024-04-03 Tanvir Mahmud , Saeed Amizadeh , Kazuhito Koishida , Diana Marculescu

Expressive speech synthesis requires vibrant prosody and well-timed pauses. We propose an effective strategy to augment a small dataset to train an expressive end-to-end Text-to-Speech model. We merge audios of emotionally congruent text…

Sound · Computer Science 2026-02-12 Raymond Chung

Task-oriented dialogue systems often assist users with personal or confidential matters. For this reason, the developers of such a system are generally prohibited from observing actual usage. So how can they know where the system is failing…

Computation and Language · Computer Science 2023-06-12 Fatemehsadat Mireshghallah , Yu Su , Tatsunori Hashimoto , Jason Eisner , Richard Shin

We propose Universal target audio Separation (UniSep), addressing the separation task on arbitrary mixtures of different types of audio. Distinguished from previous studies, UniSep is performed on unlimited source domains and unlimited…

Current deep neural network (DNN) based speech separation faces a fundamental challenge -- while the models need to be trained on short segments due to computational constraints, real-world applications typically require processing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-04 Yuzhu Wang , Archontis Politis , Konstantinos Drossos , Tuomas Virtanen

In this paper, we carry out an analysis on the use of speech separation guided diarization (SSGD) in telephone conversations. SSGD performs diarization by separating the speakers signals and then applying voice activity detection on each…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-28 Giovanni Morrone , Samuele Cornell , Desh Raj , Luca Serafini , Enrico Zovato , Alessio Brutti , Stefano Squartini

In self-supervised learning, it is challenging to reduce the gap between the enhancement performance on the estimated and target speech signals with existed pre-tasks. In this paper, we propose a multi-task pre-training method to improve…

Sound · Computer Science 2022-01-02 Yi Li , Yang Sun , Syed Mohsen Naqvi

Single-channel speech separation is a crucial task for enhancing speech recognition systems in multi-speaker environments. This paper investigates the robustness of state-of-the-art Neural Network models in scenarios where the pitch…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-23 Bunlong Lay , Sebastian Zaczek , Kristina Tesch , Timo Gerkmann

Sharing real-world speech utterances is key to the training and deployment of voice-based services. However, it also raises privacy risks as speech contains a wealth of personal data. Speaker anonymization aims to remove speaker information…

The integration of Differential Privacy (DP) with diffusion models (DMs) presents a promising yet challenging frontier, particularly due to the substantial memorization capabilities of DMs that pose significant privacy risks. Differential…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Yu-Lin Tsai , Yizhe Li , Zekai Chen , Po-Yu Chen , Chia-Mu Yu , Xuebin Ren , Francois Buet-Golfouse

We introduce PodcastMix, a dataset formalizing the task of separating background music and foreground speech in podcasts. We aim at defining a benchmark suitable for training and evaluating (deep learning) source separation models. To that…

Sound · Computer Science 2022-07-18 Nicolás Schmidt , Jordi Pons , Marius Miron

Source separation can improve automatic speech recognition (ASR) under multi-party meeting scenarios by extracting single-speaker signals from overlapped speech. Despite the success of self-supervised learning models in single-channel…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-04 Yuang Li , Xianrui Zheng , Philip C. Woodland

A great deal of recent research effort on speech spoofing countermeasures has been invested into back-end neural networks and training criteria. We contribute to this effort with a comparative perspective in this study. Our comparison of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-15 Xin Wang , Junich Yamagishi