English
Related papers

Related papers: Plug-and-Steer: Decoupling Separation and Selectio…

200 papers

This paper presents a novel approach to target speaker extraction (TSE) using Curriculum Learning (CL) techniques, addressing the challenge of distinguishing a target speaker's voice from a mixture containing interfering speakers. For…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-13 Yun Liu , Xuechen Liu , Xiaoxiao Miao , Junichi Yamagishi

This paper describes our audio-quality-based multi-strategy approach for the audio-visual target speaker extraction (AVTSE) task in the Multi-modal Information based Speech Processing (MISP) 2023 Challenge. Specifically, our approach adopts…

Sound · Computer Science 2024-03-08 Runduo Han , Xiaopeng Yan , Weiming Xu , Pengcheng Guo , Jiayao Sun , He Wang , Quan Lu , Ning Jiang , Lei Xie

Adding visual cues to audio-based speech separation can improve separation performance. This paper introduces AV-CrossNet, an audiovisual (AV) system for speech enhancement, target speaker extraction, and multi-talker speaker separation.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-18 Vahid Ahmadi Kalkhorani , Cheng Yu , Anurag Kumar , Ke Tan , Buye Xu , DeLiang Wang

Personalised speech enhancement (PSE), which extracts only the speech of a target user and removes everything else from a recorded audio clip, can potentially improve users' experiences of audio AI modules deployed in the wild. To support a…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-09 Shucong Zhang , Malcolm Chadwick , Alberto Gil C. P. Ramos , Sourav Bhattacharya

Separating target speech from mixed signals containing flexible speaker quantities presents a challenging task. While existing methods demonstrate strong separation performance and noise robustness, they predominantly assume prior knowledge…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-18 Daning Zhang , Ying Wei

Human auditory cortex excels at selectively suppressing background noise to focus on a target speaker. The process of selective attention in the brain is known to contextually exploit the available audio and visual cues to better focus on…

Sound · Computer Science 2018-09-12 Mandar Gogate , Ahsan Adeel , Ricard Marxer , Jon Barker , Amir Hussain

Speaker extraction (SE) aims to segregate the speech of a target speaker from a mixture of interfering speakers with the help of auxiliary information. Several forms of auxiliary information have been employed in single-channel SE, such as…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-18 Mohamed Elminshawi , Wolfgang Mack , Srikanth Raj Chetupalli , Soumitro Chakrabarty , Emanuël A. P. Habets

While many text-to-audio systems produce monophonic or fixed-stereo outputs, generating audio with user-defined spatial properties remains a challenge. Existing deep learning-based spatialization methods often rely on latent-space…

Sound · Computer Science 2025-09-16 Tutti Chi , Letian Gao , Yixiao Zhang

This report describes our systems submitted for the DCASE2024 Task 3 challenge: Audio and Audiovisual Sound Event Localization and Detection with Source Distance Estimation (Track B). Our main model is based on the audio-visual (AV)…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-30 Davide Berghi , Philip J. B. Jackson

This paper introduces an area-based source separation method designed for virtual meeting scenarios. The aim is to preserve speech signals from an unspecified number of sources within a defined spatial area in front of a linear microphone…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-20 Martin Strauss , Okan Köpüklü

Active speaker detection (ASD) in egocentric videos presents unique challenges due to unstable viewpoints, motion blur, and off-screen speech sources - conditions under which traditional visual-centric methods degrade significantly. We…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Yu Wang , Juhyung Ha , David J. Crandall

This paper introduces a multi-stage self-directed framework designed to address the spatial semantic segmentation of sound scene (S5) task in the DCASE 2025 Task 4 challenge. This framework integrates models focused on three distinct tasks:…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-18 Younghoo Kwon , Dongheon Lee , Dohwan Kim , Jung-Woo Choi

This paper targets a new scenario that integrates speech separation with speech compression, aiming to disentangle multiple speakers while producing discrete representations for efficient transmission or storage, with applications in online…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Hui-Peng Du , Yang Ai , Xiao-Hang Jiang , Rui-Chen Zheng , Zhen-Hua Ling

The objective of Audio-Visual Segmentation (AVS) is to localise the sounding objects within visual scenes by accurately predicting pixel-wise segmentation masks. To tackle the task, it involves a comprehensive consideration of both the data…

Computer Vision and Pattern Recognition · Computer Science 2023-10-10 Jinxiang Liu , Yu Wang , Chen Ju , Chaofan Ma , Ya Zhang , Weidi Xie

Universal sound separation (USS) aims to extract arbitrary types of sounds from real-world recordings. This can be achieved by language-queried target sound extraction (TSE), which typically consists of two components: a query network that…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-24 Hao Ma , Zhiyuan Peng , Xu Li , Mingjie Shao , Xixin Wu , Ju Liu

We propose a new problem called audio-visual segmentation (AVS), in which the goal is to output a pixel-level map of the object(s) that produce sound at the time of the image frame. To facilitate this research, we construct the first…

Computer Vision and Pattern Recognition · Computer Science 2023-01-31 Jinxing Zhou , Xuyang Shen , Jianyuan Wang , Jiayi Zhang , Weixuan Sun , Jing Zhang , Stan Birchfield , Dan Guo , Lingpeng Kong , Meng Wang , Yiran Zhong

Target speaker extraction (TSE) extracts the target speaker's voice from overlapping speech mixtures given a reference utterance. Existing approaches typically fall into two categories: discriminative and generative. Discriminative methods…

Sound · Computer Science 2026-03-16 Junwon Moon , Hyunjin Choi , Hansol Park , Heeseung Kim , Kyuhong Shim

Target speech extraction (TSE) extracts the speech of a target speaker in a mixture given auxiliary clues characterizing the speaker, such as an enrollment utterance. TSE addresses thus the challenging problem of simultaneously performing…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-15 Marc Delcroix , Keisuke Kinoshita , Tsubasa Ochiai , Katerina Zmolikova , Hiroshi Sato , Tomohiro Nakatani

Audio-Visual Segmentation (AVS) aims to achieve pixel-level localization of sound sources in videos, while Audio-Visual Semantic Segmentation (AVSS), as an extension of AVS, further pursues semantic understanding of audio-visual scenes.…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Juncheng Ma , Peiwen Sun , Yaoting Wang , Di Hu

This paper proposes a multi-task learning network with phoneme-aware and channel-wise attentive learning strategies for text-dependent Speaker Verification (SV). In the proposed structure, the frame-level multi-task learning along with the…

Sound · Computer Science 2021-06-28 Yan Liu , Zheng Li , Lin Li , Qingyang Hong
‹ Prev 1 3 4 5 6 7 10 Next ›