中文
相关论文

相关论文: ReZero: Region-customizable Sound Extraction

200 篇论文

Target speech extraction (TSE) isolates the speech of a specific speaker from a multi-talker overlapped speech mixture. Most existing TSE models rely on discriminative methods, typically predicting a time-frequency spectrogram mask for the…

音频与语音处理 · 电气工程与系统科学 2025-05-22 Hao Ma , Rujin Chen , Xiao-Lei Zhang , Ju Liu , Xuelong Li

Generating combined visual and auditory sensory experiences is critical for the consumption of immersive content. Recent advances in neural generative models have enabled the creation of high-resolution content across multiple modalities…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Rishit Dagli , Shivesh Prakash , Robert Wu , Houman Khosravani

We present AERO, a audio super-resolution model that processes speech and music signals in the spectral domain. AERO is based on an encoder-decoder architecture with U-Net like skip connections. We optimize the model using both time and…

声音 · 计算机科学 2023-02-28 Moshe Mandel , Or Tal , Yossi Adi

Environmental sound recognition (ESR) is an emerging research topic in audio pattern recognition. Many tasks are presented to resort to computational models for ESR in real-life applications. However, current models are usually designed for…

音频与语音处理 · 电气工程与系统科学 2023-11-22 Jisheng Bai , Jianfeng Chen , Mou Wang , Muhammad Saad Ayub

Speaker extraction seeks to extract the target speech in a multi-talker scenario given an auxiliary reference. Such reference can be auditory, i.e., a pre-recorded speech, visual, i.e., lip movements, or contextual, i.e., phonetic sequence.…

计算机视觉与模式识别 · 计算机科学 2022-10-13 Junjie Li , Meng Ge , Zexu Pan , Longbiao Wang , Jianwu Dang

Recently, open-vocabulary learning has emerged to accomplish segmentation for arbitrary categories of text-based descriptions, which popularizes the segmentation system to more general-purpose application scenarios. However, existing…

计算机视觉与模式识别 · 计算机科学 2023-03-31 Jie Qin , Jie Wu , Pengxiang Yan , Ming Li , Ren Yuxi , Xuefeng Xiao , Yitong Wang , Rui Wang , Shilei Wen , Xin Pan , Xingang Wang

Spatial audio, which focuses on immersive 3D sound rendering, is widely applied in the acoustic industry. One of the key problems of current spatial audio rendering methods is the lack of personalization based on different anatomies of…

计算机视觉与模式识别 · 计算机科学 2023-01-31 Xiaoyang Huang , Yanjun Wang , Yang Liu , Bingbing Ni , Wenjun Zhang , Jinxian Liu , Teng Li

Target Sound Extraction (TSE) focuses on the problem of separating sources of interest, indicated by a user's cue, from the input mixture. Most existing solutions operate in an offline fashion and are not suited to the low-latency causal…

音频与语音处理 · 电气工程与系统科学 2024-03-22 Shrishail Baligar , Mikolaj Kegler , Bryce Irvin , Marko Stamenovic , Shawn Newsam

Audio-Visual Target Speaker Extraction (AV-TSE) aims to mimic the human ability to enhance auditory perception using visual cues. Although numerous models have been proposed recently, most of them estimate target signals by primarily…

声音 · 计算机科学 2025-04-02 Wenxuan Wu , Xueyuan Chen , Shuai Wang , Jiadong Wang , Lingwei Meng , Xixin Wu , Helen Meng , Haizhou Li

We consider the task of region-based source separation of reverberant multi-microphone recordings. We assume pre-defined spatial regions with a single active source per region. The objective is to estimate the signals from the individual…

音频与语音处理 · 电气工程与系统科学 2023-03-14 Julian Wechsler , Srikanth Raj Chetupalli , Wolfgang Mack , Emanuël A. P. Habets

A three-stage approach is proposed for speaker counting and speech separation in noisy and reverberant environments. In the spatial feature extraction, a spatial coherence matrix (SCM) is computed using whitened relative transfer functions…

音频与语音处理 · 电气工程与系统科学 2023-08-08 Yicheng Hsu , Mingsian Bai

Multi-channel target speaker extraction (MC-TSE) aims to extract a target speaker's voice from multi-speaker signals captured by multiple microphones. Existing methods often rely on auxiliary clues such as direction-of-arrival (DOA) or…

音频与语音处理 · 电气工程与系统科学 2025-10-20 Tongtao Ling , Shulin He , Pengjie Shen , Zhong-Qiu Wang

This paper addresses the issue of active speaker detection (ASD) in noisy environments and formulates a robust active speaker detection (rASD) problem. Existing ASD approaches leverage both audio and visual modalities, but non-speech sounds…

多媒体 · 计算机科学 2024-04-02 Siva Sai Nagender Vasireddy , Chenxu Zhang , Xiaohu Guo , Yapeng Tian

This paper presents a robust multi-channel speaker extraction algorithm designed to handle inaccuracies in reference information. While existing approaches often rely solely on either spatial or spectral cues to identify the target speaker,…

声音 · 计算机科学 2025-12-24 Aviad Eisenberg , Sharon Gannot , Shlomo E. Chazan

Speaker extraction is to extract a target speaker's voice from multi-talker speech. It simulates humans' cocktail party effect or the selective listening ability. The prior work mostly performs speaker extraction in frequency domain, then…

音频与语音处理 · 电气工程与系统科学 2020-05-01 Chenglin Xu , Wei Rao , Eng Siong Chng , Haizhou Li

Speaker extraction aims to extract target speech signal from a multi-talker environment with interference speakers and surrounding noise, given the target speaker's reference information. Most speaker extraction systems achieve satisfactory…

音频与语音处理 · 电气工程与系统科学 2022-08-12 Chengyun Deng , Shiqian Ma , Yi Zhang , Yongtao Sha , Hui Zhang , Hui Song , Xiangang Li

Image retrieval using spoken language cues has emerged as a promising direction in multimodal perception, yet leveraging speech in multi-speaker scenarios remains challenging. We propose a novel Target Speaker Speech-Image Retrieval task…

音频与语音处理 · 电气工程与系统科学 2025-09-12 Wenhao Yang , Jianguo Wei , Wenhuan Lu , Xinyue Song , Xianghu Yue

In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances in neural approaches for Room Impulse Response (RIR)…

Speaker extraction algorithm relies on the speech sample from the target speaker as the reference point to focus its attention. Such a reference speech is typically pre-recorded. On the other hand, the temporal synchronization between…

音频与语音处理 · 电气工程与系统科学 2021-02-11 Zexu Pan , Ruijie Tao , Chenglin Xu , Haizhou Li

Supervised learning methods have shown effectiveness in estimating spatial acoustic parameters such as time difference of arrival, direct-to-reverberant ratio and reverberation time. However, they still suffer from the simulation-to-reality…

声音 · 计算机科学 2024-09-10 Bing Yang , Xiaofei Li