English
Related papers

Related papers: Target Speech Extraction with Conditional Diffusio…

200 papers

Target speaker extraction (TSE) aims to recover the speech of a desired speaker from a mixture given a short enrollment utterance, while speech enhancement (SE) focuses on improving speech quality under noisy conditions. Most existing TSE…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-21 Bang Zeng , Beilong Tang , Wang Xiang , Ming Li

This paper presents a novel approach to target speaker extraction (TSE) using Curriculum Learning (CL) techniques, addressing the challenge of distinguishing a target speaker's voice from a mixture containing interfering speakers. For…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-13 Yun Liu , Xuechen Liu , Xiaoxiao Miao , Junichi Yamagishi

Target sound extraction (TSE) consists of isolating a desired sound from a mixture of arbitrary sounds using clues to identify it. A TSE system requires solving two problems at once, identifying the target source and extracting the target…

Target speech extraction (TSE) extracts the speech of a target speaker in a mixture given auxiliary clues characterizing the speaker, such as an enrollment utterance. TSE addresses thus the challenging problem of simultaneously performing…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-15 Marc Delcroix , Keisuke Kinoshita , Tsubasa Ochiai , Katerina Zmolikova , Hiroshi Sato , Tomohiro Nakatani

Target speaker extraction (TSE) aims to isolate individual speaker voices from complex speech environments. The effectiveness of TSE systems is often compromised when the speaker characteristics are similar to each other. Recent research…

Sound · Computer Science 2024-10-08 Yun Liu , Xuechen Liu , Junichi Yamagishi

Speech enhancement aims to improve the quality of speech signals in terms of quality and intelligibility, and speech editing refers to the process of editing the speech according to specific user needs. In this paper, we propose a Unified…

Sound · Computer Science 2023-10-03 Muqiao Yang , Chunlei Zhang , Yong Xu , Zhongweiyang Xu , Heming Wang , Bhiksha Raj , Dong Yu

We propose DiffSpEx, a generative target speaker extraction method based on score-based generative modelling through stochastic differential equations. DiffSpEx deploys a continuous-time stochastic diffusion process in the complex…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-10 Theodor Nguyen , Guangzhi Sun , Xianrui Zheng , Chao Zhang , Philip C Woodland

Target speaker extraction (TSE) aims to isolate a specific voice from multiple mixed speakers relying on a registerd sample. Since voiceprint features usually vary greatly, current end-to-end neural networks require large model parameters…

Sound · Computer Science 2024-01-08 Shulin He , Jinjiang liu , Hao Li , Yang Yang , Fei Chen , Xueliang Zhang

Diffusion probabilistic models have demonstrated an outstanding capability to model natural images and raw audio waveforms through a paired diffusion and reverse processes. The unique property of the reverse process (namely, eliminating…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-23 Yen-Ju Lu , Yu Tsao , Shinji Watanabe

Sound separation (SS) and target sound extraction (TSE) are fundamental techniques for addressing complex acoustic scenarios. While existing SS methods struggle with determining the unknown number of sound sources, TSE approaches require…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-25 Hongyu Wang , Chenda Li , Xin Zhou , Shuai Wang , Yanmin Qian

In target speaker extraction (TSE), we aim to recover target speech from a multi-talker mixture using a short enrollment utterance as reference. Recent studies on diffusion and flow-matching generators have improved target-speech fidelity.…

Sound · Computer Science 2026-03-12 Duojia Li , Shuhan Zhang , Zihan Qian , Wenxuan Wu , Shuai Wang , Qingyang Hong , Lin Li , Haizhou Li

Generative target speaker extraction (TSE) methods often produce more natural outputs than predictive models. Recent work based on diffusion or flow matching (FM) typically relies on a small, fixed number of reverse steps with a fixed step…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-21 Tsun-An Hsieh , Minje Kim

Target Speaker Extraction (TSE) plays a critical role in enhancing speech signals in noisy and multi-speaker environments. This paper presents an end-to-end TSE model that incorporates Direction of Arrival (DOA) and beamwidth embeddings to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-29 Kangqi Jing , Wenbin Zhang , Yu Gao

Target speaker extraction (TSE) extracts the target speaker's voice from overlapping speech mixtures given a reference utterance. Existing approaches typically fall into two categories: discriminative and generative. Discriminative methods…

Sound · Computer Science 2026-03-16 Junwon Moon , Hyunjin Choi , Hansol Park , Heeseung Kim , Kyuhong Shim

In this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing the previous U-Net backbone with a skip-connected Transformer…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-03 Helin Wang , Jiarui Hai , Yen-Ju Lu , Karan Thakkar , Mounya Elhilali , Najim Dehak

Target Sound Extraction (TSE) focuses on the problem of separating sources of interest, indicated by a user's cue, from the input mixture. Most existing solutions operate in an offline fashion and are not suited to the low-latency causal…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-22 Shrishail Baligar , Mikolaj Kegler , Bryce Irvin , Marko Stamenovic , Shawn Newsam

Diffusion-based speech enhancement (SE) models need to incorporate correct prior knowledge as reliable conditions to generate accurate predictions. However, providing reliable conditions using noisy features is challenging. One solution is…

Sound · Computer Science 2025-10-08 Hao Shi , Xugang Lu , Kazuki Shimada , Tatsuya Kawahara

Target speech extraction (TSE) systems are designed to extract target speech from a multi-talker mixture. The popular training objective for most prior TSE networks is to enhance reconstruction performance of extracted speech waveform.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-10 Kai Liu , Ziqing Du , Xucheng Wan , Huan Zhou

Developing a robust speech emotion recognition (SER) system in noisy conditions faces challenges posed by different noise properties. Most previous studies have not considered the impact of human speech noise, thus limiting the application…

Sound · Computer Science 2024-12-18 Jinyi Mi , Xiaohan Shi , Ding Ma , Jiajun He , Takuya Fujimura , Tomoki Toda

The goal of speech enhancement (SE) is to eliminate the background interference from the noisy speech signal. Generative models such as diffusion models (DM) have been applied to the task of SE because of better generalization in unseen…

Sound · Computer Science 2023-09-06 Wen Wang , Dongchao Yang , Qichen Ye , Bowen Cao , Yuexian Zou