中文
相关论文

相关论文: Binaural Target Speaker Extraction using Individua…

200 篇论文

This paper addresses the problem of single-channel speech separation, where the number of speakers is unknown, and each speaker may speak multiple utterances. We propose a speech separation model that simultaneously performs separation,…

音频与语音处理 · 电气工程与系统科学 2025-05-23 Yuzhu Wang , Archontis Politis , Konstantinos Drossos , Tuomas Virtanen

Personalized Head-Related Transfer Functions (HRTFs) are starting to be introduced in many commercial immersive audio applications and are crucial for realistic spatial audio rendering. However, one of the main hesitations regarding their…

声音 · 计算机科学 2025-10-03 Xuyi Hu , Jian Li , Shaojie Zhang , Stefan Goetz , Lorenzo Picinali , Ozgur B. Akan , Aidan O. T. Hogg

The development of privacy-preserving automatic speaker verification systems has been the focus of a number of studies with the intent of allowing users to authenticate themselves without risking the privacy of their voice. However, current…

音频与语音处理 · 电气工程与系统科学 2022-10-28 Francisco Teixeira , Alberto Abad , Bhiksha Raj , Isabel Trancoso

In a scenario with multiple persons talking simultaneously, the spatial characteristics of the signals are the most distinct feature for extracting the target signal. In this work, we develop a deep joint spatial-spectral non-linear filter…

音频与语音处理 · 电气工程与系统科学 2023-04-05 Kristina Tesch , Timo Gerkmann

Head Related Transfer Functions (HRTFs) play a crucial role in creating immersive spatial audio experiences. However, HRTFs differ significantly from person to person, and traditional methods for estimating personalized HRTFs are expensive,…

音频与语音处理 · 电气工程与系统科学 2023-11-08 Vivek Jayaram , Ira Kemelmacher-Shlizerman , Steven M. Seitz

In this work, we present a two-stage method for speaker extraction under reverberant and noisy conditions. Given a reference signal of the desired speaker, the clean, but the still reverberant, desired speaker is first extracted from the…

声音 · 计算机科学 2023-03-14 Aviad Eisenberg , Sharon Gannot , Shlomo E. Chazan

Speaker-aware source separation methods are promising workarounds for major difficulties such as arbitrary source permutation and unknown number of sources. However, it remains challenging to achieve satisfying performance provided a very…

声音 · 计算机科学 2018-07-25 Jun Wang , Jie Chen , Dan Su , Lianwu Chen , Meng Yu , Yanmin Qian , Dong Yu

The vast majority of speech separation methods assume that the number of speakers is known in advance, hence they are specific to the number of speakers. By contrast, a more realistic and challenging task is to separate a mixture in which…

声音 · 计算机科学 2022-03-31 Zhenhao Jin , Xiang Hao , Xiangdong Su

This paper investigates the use of relative cues for text-based target speech extraction (TSE). We first provide a theoretical justification for relative cues from the perspectives of human perception and label quantization, showing that…

音频与语音处理 · 电气工程与系统科学 2026-04-03 Wang Dai , Archontis Politis , Tuomas Virtanen

A new database of head-related transfer functions (HRTFs) for accurate sound source localization is presented through precise measurement and post-processing in terms of improved frequency bandwidth and causality of head-related impulse…

音频与语音处理 · 电气工程与系统科学 2022-04-07 Gyeong-Tae Lee , Sang-Min Choi , Byeong-Yun Ko , Yong-Hwa Park

Strong representations of target speakers can help extract important information about speakers and detect corresponding temporal regions in multi-speaker conversations. In this study, we propose a neural architecture that simultaneously…

声音 · 计算机科学 2023-06-07 Chin-Yi Cheng , Hung-Shin Lee , Yu Tsao , Hsin-Min Wang

Zero-shot multi-speaker TTS aims to synthesize speech with the voice of a chosen target speaker without any fine-tuning. Prevailing methods, however, encounter limitations at adapting to new speakers of out-of-domain settings, primarily due…

声音 · 计算机科学 2024-03-06 Yejin Jeon , Yunsu Kim , Gary Geunbae Lee

Estimating Head-Related Transfer Functions (HRTFs) of arbitrary source points is essential in immersive binaural audio rendering. Computing each individual's HRTFs is challenging, as traditional approaches require expensive time and…

音频与语音处理 · 电气工程与系统科学 2022-11-04 Jin Woo Lee , Sungho Lee , Kyogu Lee

In this paper we present a unified time-frequency method for speaker extraction in clean and noisy conditions. Given a mixed signal, along with a reference signal, the common approaches for extracting the desired speaker are either applied…

声音 · 计算机科学 2022-03-08 Aviad Eisenberg , Sharon Gannot , Shlomo E. Chazan

We present the first neural target speech extraction (TSE) system that uses human feedback for iterative refinement. Our approach allows users to mark specific segments of the TSE output, generating an edit mask. The refinement system then…

声音 · 计算机科学 2025-08-06 Malek Itani , Ashton Graves , Sefik Emre Eskimez , Shyamnath Gollakota

Speech separation involves extracting an individual speaker's voice from a multi-speaker audio signal. The increasing complexity of real-world environments, where multiple speakers might converse simultaneously, underscores the importance…

音频与语音处理 · 电气工程与系统科学 2024-01-09 Renana Opochinsky , Mordehay Moradi , Sharon Gannot

Recently, speech enhancement technologies that are based on deep learning have received considerable research attention. If the spatial information in microphone signals is exploited, microphone arrays can be advantageous under some adverse…

音频与语音处理 · 电气工程与系统科学 2022-07-19 Yicheng Hsu , Yonghan Lee , Mingsian R. Bai

Personalized head-related transfer functions (HRTFs) are essential for ensuring a realistic auditory experience over headphones, because they take into account individual anatomical differences that affect listening. Most machine learning…

音频与语音处理 · 电气工程与系统科学 2026-01-27 You Zhang , Andrew Francl , Ruohan Gao , Paul Calamia , Zhiyao Duan , Ishwarya Ananthabhotla

This paper addresses the problem of speech separation and enhancement from multichannel convolutive and noisy mixtures, \emph{assuming known mixing filters}. We propose to perform the speech separation and enhancement task in the short-time…

声音 · 计算机科学 2019-01-31 Xiaofei Li , Laurent Girin , Sharon Gannot , Radu Horaud

State-of-the-art target speaker extraction (TSE) systems are typically designed to generalize to any given mixing environment, necessitating a model with a large enough capacity as a generalist. Personalized speech enhancement could be a…

音频与语音处理 · 电气工程与系统科学 2025-08-06 Tsun-An Hsieh , Minje Kim