中文
相关论文

相关论文: Plug-and-Steer: Decoupling Separation and Selectio…

200 篇论文

This paper presents a novel approach to target speaker extraction (TSE) using Curriculum Learning (CL) techniques, addressing the challenge of distinguishing a target speaker's voice from a mixture containing interfering speakers. For…

音频与语音处理 · 电气工程与系统科学 2024-06-13 Yun Liu , Xuechen Liu , Xiaoxiao Miao , Junichi Yamagishi

This paper describes our audio-quality-based multi-strategy approach for the audio-visual target speaker extraction (AVTSE) task in the Multi-modal Information based Speech Processing (MISP) 2023 Challenge. Specifically, our approach adopts…

声音 · 计算机科学 2024-03-08 Runduo Han , Xiaopeng Yan , Weiming Xu , Pengcheng Guo , Jiayao Sun , He Wang , Quan Lu , Ning Jiang , Lei Xie

Adding visual cues to audio-based speech separation can improve separation performance. This paper introduces AV-CrossNet, an audiovisual (AV) system for speech enhancement, target speaker extraction, and multi-talker speaker separation.…

音频与语音处理 · 电气工程与系统科学 2024-06-18 Vahid Ahmadi Kalkhorani , Cheng Yu , Anurag Kumar , Ke Tan , Buye Xu , DeLiang Wang

Personalised speech enhancement (PSE), which extracts only the speech of a target user and removes everything else from a recorded audio clip, can potentially improve users' experiences of audio AI modules deployed in the wild. To support a…

音频与语音处理 · 电气工程与系统科学 2022-11-09 Shucong Zhang , Malcolm Chadwick , Alberto Gil C. P. Ramos , Sourav Bhattacharya

Separating target speech from mixed signals containing flexible speaker quantities presents a challenging task. While existing methods demonstrate strong separation performance and noise robustness, they predominantly assume prior knowledge…

音频与语音处理 · 电气工程与系统科学 2025-07-18 Daning Zhang , Ying Wei

Human auditory cortex excels at selectively suppressing background noise to focus on a target speaker. The process of selective attention in the brain is known to contextually exploit the available audio and visual cues to better focus on…

声音 · 计算机科学 2018-09-12 Mandar Gogate , Ahsan Adeel , Ricard Marxer , Jon Barker , Amir Hussain

Speaker extraction (SE) aims to segregate the speech of a target speaker from a mixture of interfering speakers with the help of auxiliary information. Several forms of auxiliary information have been employed in single-channel SE, such as…

音频与语音处理 · 电气工程与系统科学 2023-09-18 Mohamed Elminshawi , Wolfgang Mack , Srikanth Raj Chetupalli , Soumitro Chakrabarty , Emanuël A. P. Habets

While many text-to-audio systems produce monophonic or fixed-stereo outputs, generating audio with user-defined spatial properties remains a challenge. Existing deep learning-based spatialization methods often rely on latent-space…

声音 · 计算机科学 2025-09-16 Tutti Chi , Letian Gao , Yixiao Zhang

This report describes our systems submitted for the DCASE2024 Task 3 challenge: Audio and Audiovisual Sound Event Localization and Detection with Source Distance Estimation (Track B). Our main model is based on the audio-visual (AV)…

音频与语音处理 · 电气工程与系统科学 2024-10-30 Davide Berghi , Philip J. B. Jackson

This paper introduces an area-based source separation method designed for virtual meeting scenarios. The aim is to preserve speech signals from an unspecified number of sources within a defined spatial area in front of a linear microphone…

音频与语音处理 · 电气工程与系统科学 2024-08-20 Martin Strauss , Okan Köpüklü

Active speaker detection (ASD) in egocentric videos presents unique challenges due to unstable viewpoints, motion blur, and off-screen speech sources - conditions under which traditional visual-centric methods degrade significantly. We…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Yu Wang , Juhyung Ha , David J. Crandall

This paper introduces a multi-stage self-directed framework designed to address the spatial semantic segmentation of sound scene (S5) task in the DCASE 2025 Task 4 challenge. This framework integrates models focused on three distinct tasks:…

音频与语音处理 · 电气工程与系统科学 2025-09-18 Younghoo Kwon , Dongheon Lee , Dohwan Kim , Jung-Woo Choi

This paper targets a new scenario that integrates speech separation with speech compression, aiming to disentangle multiple speakers while producing discrete representations for efficient transmission or storage, with applications in online…

音频与语音处理 · 电气工程与系统科学 2026-01-21 Hui-Peng Du , Yang Ai , Xiao-Hang Jiang , Rui-Chen Zheng , Zhen-Hua Ling

The objective of Audio-Visual Segmentation (AVS) is to localise the sounding objects within visual scenes by accurately predicting pixel-wise segmentation masks. To tackle the task, it involves a comprehensive consideration of both the data…

计算机视觉与模式识别 · 计算机科学 2023-10-10 Jinxiang Liu , Yu Wang , Chen Ju , Chaofan Ma , Ya Zhang , Weidi Xie

Universal sound separation (USS) aims to extract arbitrary types of sounds from real-world recordings. This can be achieved by language-queried target sound extraction (TSE), which typically consists of two components: a query network that…

音频与语音处理 · 电气工程与系统科学 2025-03-24 Hao Ma , Zhiyuan Peng , Xu Li , Mingjie Shao , Xixin Wu , Ju Liu

We propose a new problem called audio-visual segmentation (AVS), in which the goal is to output a pixel-level map of the object(s) that produce sound at the time of the image frame. To facilitate this research, we construct the first…

计算机视觉与模式识别 · 计算机科学 2023-01-31 Jinxing Zhou , Xuyang Shen , Jianyuan Wang , Jiayi Zhang , Weixuan Sun , Jing Zhang , Stan Birchfield , Dan Guo , Lingpeng Kong , Meng Wang , Yiran Zhong

Target speaker extraction (TSE) extracts the target speaker's voice from overlapping speech mixtures given a reference utterance. Existing approaches typically fall into two categories: discriminative and generative. Discriminative methods…

声音 · 计算机科学 2026-03-16 Junwon Moon , Hyunjin Choi , Hansol Park , Heeseung Kim , Kyuhong Shim

Target speech extraction (TSE) extracts the speech of a target speaker in a mixture given auxiliary clues characterizing the speaker, such as an enrollment utterance. TSE addresses thus the challenging problem of simultaneously performing…

音频与语音处理 · 电气工程与系统科学 2022-07-15 Marc Delcroix , Keisuke Kinoshita , Tsubasa Ochiai , Katerina Zmolikova , Hiroshi Sato , Tomohiro Nakatani

Audio-Visual Segmentation (AVS) aims to achieve pixel-level localization of sound sources in videos, while Audio-Visual Semantic Segmentation (AVSS), as an extension of AVS, further pursues semantic understanding of audio-visual scenes.…

计算机视觉与模式识别 · 计算机科学 2024-09-13 Juncheng Ma , Peiwen Sun , Yaoting Wang , Di Hu

This paper proposes a multi-task learning network with phoneme-aware and channel-wise attentive learning strategies for text-dependent Speaker Verification (SV). In the proposed structure, the frame-level multi-task learning along with the…

声音 · 计算机科学 2021-06-28 Yan Liu , Zheng Li , Lin Li , Qingyang Hong