English
Related papers

Related papers: Improving curriculum learning for target speaker e…

200 papers

Target sound extraction (TSE) separates the target sound from the mixture signals based on provided clues. However, the performance of existing models significantly degrades under reverberant conditions. Inspired by auditory scene analysis…

Sound · Computer Science 2024-06-14 Yiwen Wang , Xihong Wu

Target speaker extraction (TSE) aims to isolate a desired speaker's voice from a multi-speaker mixture using auxiliary information such as a reference utterance. Although recent advances in diffusion and flow-matching models have improved…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-23 Riki Shimizu , Xilin Jiang , Nima Mesgarani

In this paper, we investigate a novel approach for Target Speech Extraction (TSE), which relies solely on textual context to extract the target speech. We refer to this task as Contextual Speech Extraction (CSE). Unlike traditional TSE…

Sound · Computer Science 2025-03-13 Minsu Kim , Rodrigo Mira , Honglie Chen , Stavros Petridis , Maja Pantic

Sound separation (SS) and target sound extraction (TSE) are fundamental techniques for addressing complex acoustic scenarios. While existing SS methods struggle with determining the unknown number of sound sources, TSE approaches require…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-25 Hongyu Wang , Chenda Li , Xin Zhou , Shuai Wang , Yanmin Qian

Developing a robust speech emotion recognition (SER) system in noisy conditions faces challenges posed by different noise properties. Most previous studies have not considered the impact of human speech noise, thus limiting the application…

Sound · Computer Science 2024-12-18 Jinyi Mi , Xiaohan Shi , Ding Ma , Jiajun He , Takuya Fujimura , Tomoki Toda

Personalized speech enhancement (PSE) models can improve the audio quality of teleconferencing systems by adapting to the characteristics of a speaker's voice. However, most existing methods require a separate speaker embedding model to…

Sound · Computer Science 2024-06-17 Tanel Pärnamaa , Ando Saabas

Targeted syntactic evaluation of subject-verb number agreement in English (TSE) evaluates language models' syntactic knowledge using hand-crafted minimal pairs of sentences that differ only in the main verb's conjugation. The method…

Computation and Language · Computer Science 2021-04-21 Benjamin Newman , Kai-Siang Ang , Julia Gong , John Hewitt

Target speaker extraction (TSE) focuses on isolating the speech of a specific target speaker from overlapped multi-talker speech, which is a typical setup in the cocktail party problem. In recent years, TSE draws increasing attention due to…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-25 Shuai Wang , Ke Zhang , Shaoxiong Lin , Junjie Li , Xuefei Wang , Meng Ge , Jianwei Yu , Yanmin Qian , Haizhou Li

This paper aims to achieve single-channel target speech extraction (TSE) in enclosures by solely utilizing distance information. This is the first work that utilizes only distance cues without using speaker physiological information for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-31 Runwu Shi , Benjamin Yen , Kazuhiro Nakadai

Target Speech Extraction (TSE) aims to isolate a target speaker's voice from a mixture of multiple speakers by leveraging speaker-specific cues, typically provided as auxiliary audio (a.k.a. cue audio). Although recent advancements in TSE…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-09 Helin Wang , Jiarui Hai , Dongchao Yang , Chen Chen , Kai Li , Junyi Peng , Thomas Thebaud , Laureano Moro Velazquez , Jesus Villalba , Najim Dehak

We propose listen to extract (LExt), a highly-effective while extremely-simple algorithm for monaural target speaker extraction (TSE). Given an enrollment utterance of a target speaker, LExt aims at extracting the target speaker from the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-06 Pengjie Shen , Kangrui Chen , Shulin He , Pengru Chen , Shuqi Yuan , He Kong , Xueliang Zhang , Zhong-Qiu Wang

In cross-lingual speech synthesis, the speech in various languages can be synthesized for a monoglot speaker. Normally, only the data of monoglot speakers are available for model training, thus the speaker similarity is relatively low…

Sound · Computer Science 2022-01-21 J. Yang , Lei He

Personalized speech enhancement (PSE) is a real-time SE approach utilizing a speaker embedding of a target person to remove background noise, reverberation, and interfering voices. To deploy a PSE model for full duplex communications, the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-29 Sefik Emre Eskimez , Takuya Yoshioka , Alex Ju , Min Tang , Tanel Parnamaa , Huaming Wang

Audio-visual target speaker extraction (AV-TSE) models primarily rely on visual cues from the target speaker. However, humans also leverage linguistic knowledge, such as syntactic constraints, next word prediction, and prior knowledge of…

Sound · Computer Science 2025-11-11 Wenxuan Wu , Shuai Wang , Xixin Wu , Helen Meng , Haizhou Li

Audio-visual target speaker extraction (AV-TSE) aims to extract the specific person's speech from the audio mixture given auxiliary visual cues. Previous methods usually search for the target voice through speech-lip synchronization.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-04 Ruijie Tao , Xinyuan Qian , Yidi Jiang , Junjie Li , Jiadong Wang , Haizhou Li

This paper investigates the use of relative cues for text-based target speech extraction (TSE). We first provide a theoretical justification for relative cues from the perspectives of human perception and label quantization, showing that…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-03 Wang Dai , Archontis Politis , Tuomas Virtanen

Deep learning models are becoming predominant in many fields of machine learning. Text-to-Speech (TTS), the process of synthesizing artificial speech from text, is no exception. To this end, a deep neural network is usually trained using a…

Sound · Computer Science 2021-02-11 Giuseppe Ruggiero , Enrico Zovato , Luigi Di Caro , Vincent Pollet

Target confusion, defined as occasional switching to non-target speakers, poses a key challenge for end-to-end speaker extraction (E2E-SE) systems. We argue that this problem is largely caused by the lack of generalizability and…

Sound · Computer Science 2025-05-29 Zhenghai You , Zhenyu Zhou , Lantian Li , Dong Wang

Target speaker extraction aims to separate the voice of a specific speaker from mixed speech. Traditionally, this process has relied on extracting a speaker embedding from a reference speech, in which a speaker recognition model is…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-21 Bang Zeng , Ming Li

The goal of text-queried target sound extraction (TSE) is to extract from a mixture a sound source specified with a natural-language caption. While it is preferable to have access to large-scale text-audio pairs to address a variety of text…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-23 Kohei Saijo , Janek Ebbers , François G. Germain , Sameer Khurana , Gordon Wichern , Jonathan Le Roux