English
Related papers

Related papers: VorTEX: Various overlap ratio for Target speech EX…

200 papers

Target speaker extraction (TSE) extracts the target speaker's voice from overlapping speech mixtures given a reference utterance. Existing approaches typically fall into two categories: discriminative and generative. Discriminative methods…

Sound · Computer Science 2026-03-16 Junwon Moon , Hyunjin Choi , Hansol Park , Heeseung Kim , Kyuhong Shim

Speaker extraction (SE) aims to segregate the speech of a target speaker from a mixture of interfering speakers with the help of auxiliary information. Several forms of auxiliary information have been employed in single-channel SE, such as…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-18 Mohamed Elminshawi , Wolfgang Mack , Srikanth Raj Chetupalli , Soumitro Chakrabarty , Emanuël A. P. Habets

Real-time target speaker extraction (TSE) is intended to extract the desired speaker's voice from the observed mixture of multiple speakers in a streaming manner. Implementing real-time TSE is challenging as the computational complexity…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-03 Hiroshi Sato , Takafumi Moriya , Masato Mimura , Shota Horiguchi , Tsubasa Ochiai , Takanori Ashihara , Atsushi Ando , Kentaro Shinayama , Marc Delcroix

Audio-visual target speaker extraction (AV-TSE) models primarily rely on visual cues from the target speaker. However, humans also leverage linguistic knowledge, such as syntactic constraints, next word prediction, and prior knowledge of…

Sound · Computer Science 2025-11-11 Wenxuan Wu , Shuai Wang , Xixin Wu , Helen Meng , Haizhou Li

Target sound extraction (TSE) separates the target sound from the mixture signals based on provided clues. However, the performance of existing models significantly degrades under reverberant conditions. Inspired by auditory scene analysis…

Sound · Computer Science 2024-06-14 Yiwen Wang , Xihong Wu

Target speaker extraction (TSE) aims to isolate a desired speaker's voice from a multi-speaker mixture using auxiliary information such as a reference utterance. Although recent advances in diffusion and flow-matching models have improved…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-23 Riki Shimizu , Xilin Jiang , Nima Mesgarani

Personalized or target speech extraction (TSE) typically needs a clean enrollment -- hard to obtain in real-world crowded environments. We remove the essential need for enrollment by predicting, from the mixture itself, a small set of…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-06 FNU Sidharth , Meysam Asgari , Hao-Wen Dong , Dhruv Jain

Target speaker extraction (TSE) aims to extract the target speaker's voice from the input mixture. Previous studies have concentrated on high-overlapping scenarios. However, real-world applications usually meet more complex scenarios like…

Sound · Computer Science 2024-01-30 He Zhao , Hangting Chen , Jianwei Yu , Yuehai Wang

Audio-visual target speaker extraction (AV-TSE) models primarily rely on target visual cues to isolate the target speaker's voice from others. We know that humans leverage linguistic knowledge, such as syntax and semantics, to support…

Sound · Computer Science 2025-06-17 Wenxuan Wu , Shuai Wang , Xixin Wu , Helen Meng , Haizhou Li

Target speaker extraction aims to extract the speech of a specific speaker from a multi-talker mixture as specified by an auxiliary reference. Most studies focus on the scenario where the target speech is highly overlapped with the…

Sound · Computer Science 2023-09-18 Junjie Li , Ruijie Tao , Zexu Pan , Meng Ge , Shuai Wang , Haizhou Li

This paper describes an audio-visual speech enhancement (AV-SE) method that estimates from noisy input audio a mixture of the speech of the speaker appearing in an input video (on-screen target speech) and of a selected speaker not…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-13 Tomoya Yoshinaga , Keitaro Tanaka , Shigeo Morishima

Research on audio clue-based target speaker extraction (TSE) has focused on modeling mixtures and reference speech, achieving strong results in English due to abundant datasets. However, cross-lingual properties remain underexplored, as…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-10 The Hieu Pham , Phuong Thanh Tran Nguyen , Xuan Tho Nguyen , Tan Dat Nguyen , Duc Dung Nguyen

Target Speech Extraction (TSE) aims to isolate a target speaker's voice from a mixture of multiple speakers by leveraging speaker-specific cues, typically provided as auxiliary audio (a.k.a. cue audio). Although recent advancements in TSE…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-09 Helin Wang , Jiarui Hai , Dongchao Yang , Chen Chen , Kai Li , Junyi Peng , Thomas Thebaud , Laureano Moro Velazquez , Jesus Villalba , Najim Dehak

In this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing the previous U-Net backbone with a skip-connected Transformer…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-03 Helin Wang , Jiarui Hai , Yen-Ju Lu , Karan Thakkar , Mounya Elhilali , Najim Dehak

Sound separation (SS) and target sound extraction (TSE) are fundamental techniques for addressing complex acoustic scenarios. While existing SS methods struggle with determining the unknown number of sound sources, TSE approaches require…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-25 Hongyu Wang , Chenda Li , Xin Zhou , Shuai Wang , Yanmin Qian

Personalized speech enhancement (PSE) models utilize additional cues, such as speaker embeddings like d-vectors, to remove background noise and interfering speech in real-time and thus improve the speech quality of online video conferencing…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-20 Sefik Emre Eskimez , Takuya Yoshioka , Huaming Wang , Xiaofei Wang , Zhuo Chen , Xuedong Huang

The primary goal of multi-modal TSE (MTSE) is to extract a target speaker from a speech mixture using complementary information from different modalities, such as audio enrolment and visual feeds corresponding to the target speaker. MTSE…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-10 Srikanth Korse , Mohamed Elminshawi , Emanuel A. P. Habets , Srikanth Raj Chetupalli

This paper describes our audio-quality-based multi-strategy approach for the audio-visual target speaker extraction (AVTSE) task in the Multi-modal Information based Speech Processing (MISP) 2023 Challenge. Specifically, our approach adopts…

Sound · Computer Science 2024-03-08 Runduo Han , Xiaopeng Yan , Weiming Xu , Pengcheng Guo , Jiayao Sun , He Wang , Quan Lu , Ning Jiang , Lei Xie

A speaker extraction algorithm seeks to extract the target speaker's speech from a multi-talker speech mixture. The prior studies focus mostly on speaker extraction from a highly overlapped multi-talker speech mixture. However, the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-01 Zexu Pan , Meng Ge , Haizhou Li

Target speaker extraction focuses on isolating a specific speaker's voice from an audio mixture containing multiple speakers. To provide information about the target speaker's identity, prior works have utilized clean audio samples as…

Sound · Computer Science 2025-12-09 Shitong Xu , Yiyuan Yang , Niki Trigoni , Andrew Markham