English
Related papers

Related papers: HRTF-guided Binaural Target Speaker Extraction wit…

200 papers

Audio-Visual Target Speaker Extraction (AV-TSE) aims to mimic the human ability to enhance auditory perception using visual cues. Although numerous models have been proposed recently, most of them estimate target signals by primarily…

Sound · Computer Science 2025-04-02 Wenxuan Wu , Xueyuan Chen , Shuai Wang , Jiadong Wang , Lingwei Meng , Xixin Wu , Helen Meng , Haizhou Li

Target speaker extraction (TSE) extracts the target speaker's voice from overlapping speech mixtures given a reference utterance. Existing approaches typically fall into two categories: discriminative and generative. Discriminative methods…

Sound · Computer Science 2026-03-16 Junwon Moon , Hyunjin Choi , Hansol Park , Heeseung Kim , Kyuhong Shim

Target sound extraction (TSE) consists of isolating a desired sound from a mixture of arbitrary sounds using clues to identify it. A TSE system requires solving two problems at once, identifying the target source and extracting the target…

The remarkable ability of humans to selectively focus on a target speaker in cocktail party scenarios is facilitated by binaural audio processing. In this paper, we present a binaural time-domain Target Speaker Extraction model based on the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-19 Hanyu Meng , Qiquan Zhang , Xiangyu Zhang , Vidhyasaharan Sethu , Eliathamby Ambikairajah

In hearing aid applications, an important objective is to accurately estimate the direction of arrival (DOA) of multiple speakers in noisy and reverberant environments. Recently, we proposed a binaural DOA estimation method, where the DOAs…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-11 Daniel Fejgin , Simon Doclo

In target speaker extraction (TSE), we aim to recover target speech from a multi-talker mixture using a short enrollment utterance as reference. Recent studies on diffusion and flow-matching generators have improved target-speech fidelity.…

Sound · Computer Science 2026-03-12 Duojia Li , Shuhan Zhang , Zihan Qian , Wenxuan Wu , Shuai Wang , Qingyang Hong , Lin Li , Haizhou Li

Audio-visual target speaker extraction (AV-TSE) models primarily rely on visual cues from the target speaker. However, humans also leverage linguistic knowledge, such as syntactic constraints, next word prediction, and prior knowledge of…

Sound · Computer Science 2025-11-11 Wenxuan Wu , Shuai Wang , Xixin Wu , Helen Meng , Haizhou Li

Target sound extraction (TSE) separates the target sound from the mixture signals based on provided clues. However, the performance of existing models significantly degrades under reverberant conditions. Inspired by auditory scene analysis…

Sound · Computer Science 2024-06-14 Yiwen Wang , Xihong Wu

Speech clarity and spatial audio immersion are the two most critical factors in enhancing remote conferencing experiences. Existing methods are often limited: either due to the lack of spatial information when using only one microphone, or…

Sound · Computer Science 2025-07-14 Cheng Chi , Xiaoyu Li , Yuxuan Ke , Qunping Ni , Yao Ge , Xiaodong Li , Chengshi Zheng

Generative target speaker extraction (TSE) methods often produce more natural outputs than predictive models. Recent work based on diffusion or flow matching (FM) typically relies on a small, fixed number of reverse steps with a fixed step…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-21 Tsun-An Hsieh , Minje Kim

Target speaker extraction aims to isolate a specific speaker's voice from a composite of multiple sound sources, guided by an enrollment utterance or called anchor. Current methods predominantly derive speaker embeddings from the anchor and…

Sound · Computer Science 2024-01-08 Shulin He , Huaiwen Zhang , Wei Rao , Kanghao Zhang , Yukai Ju , Yang Yang , Xueliang Zhang

Head-related transfer functions (HRTFs) with dense spatial grids are desired for immersive binaural audio generation, but their recording is time-consuming. Although HRTF spatial upsampling has shown remarkable progress with neural fields,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-23 Yoshiki Masuyama , Gordon Wichern , François G. Germain , Christopher Ick , Jonathan Le Roux

Target sound extraction (TSE) aims to extract the sound part of a target sound event class from a mixture audio with multiple sound events. The previous works mainly focus on the problems of weakly-labelled data, jointly learning and new…

Sound · Computer Science 2022-04-05 Helin Wang , Dongchao Yang , Chao Weng , Jianwei Yu , Yuexian Zou

Target speaker extraction (TSE) aims to isolate individual speaker voices from complex speech environments. The effectiveness of TSE systems is often compromised when the speaker characteristics are similar to each other. Recent research…

Sound · Computer Science 2024-10-08 Yun Liu , Xuechen Liu , Junichi Yamagishi

We present a head-related transfer function (HRTF) estimation method which relies on a data-driven prior given by a score-based diffusion model. The HRTF is estimated in reverberant environments using natural excitation signals, e.g. human…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-03 Etienne Thuillier , Jean-Marie Lemercier , Eloi Moliner , Timo Gerkmann , Vesa Välimäki

In target speaker extraction, many studies rely on the speaker embedding which is obtained from an enrollment of the target speaker and employed as the guidance. However, solely using speaker embedding may not fully utilize the contextual…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-28 Xue Yang , Changchun Bao , Jing Zhou , Xianhong Chen

Audio-Visual Target Speaker Extraction (AVTSE) aims to isolate a target speaker's voice in a multi-speaker environment with visual cues as auxiliary. Most of the existing AVTSE methods encode visual and audio features simultaneously,…

Sound · Computer Science 2025-11-13 Zixuan Li , Xueliang Zhang , Lei Miao , Zhipeng Yan , Ying Sun , Chong Zhu

We propose a method of head-related transfer function (HRTF) interpolation from sparsely measured HRTFs using an autoencoder with source position conditioning. The proposed method is drawn from an analogy between an HRTF interpolation…

Sound · Computer Science 2022-07-25 Yuki Ito , Tomohiko Nakamura , Shoichi Koyama , Hiroshi Saruwatari

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world scenarios, visual cues are not always available due to various…

Sound · Computer Science 2025-04-01 Junjie Li , Ke Zhang , Shuai Wang , Kong Aik Lee , Man-Wai Mak , Haizhou Li

Headphone-based spatial audio uses head-related transfer functions (HRTFs) to simulate real-world acoustic environments. HRTFs are unique to everyone, due to personal morphology, shaping how sound waves interact with the body before…

‹ Prev 1 3 4 5 6 7 10 Next ›