English
Related papers

Related papers: MeanFlow-TSE: One-Step Generative Target Speaker E…

200 papers

The goal of text-queried target sound extraction (TSE) is to extract from a mixture a sound source specified with a natural-language caption. While it is preferable to have access to large-scale text-audio pairs to address a variety of text…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-23 Kohei Saijo , Janek Ebbers , François G. Germain , Sameer Khurana , Gordon Wichern , Jonathan Le Roux

Target sound extraction (TSE) separates the target sound from the mixture signals based on provided clues. However, the performance of existing models significantly degrades under reverberant conditions. Inspired by auditory scene analysis…

Sound · Computer Science 2024-06-14 Yiwen Wang , Xihong Wu

Recent years have witnessed remarkable progress in Text-to-Audio Generation (TTA), providing sound creators with powerful tools to transform inspirations into vivid audio. Yet despite these advances, current TTA systems often suffer from…

Sound · Computer Science 2025-10-23 Xiquan Li , Junxi Liu , Yuzhe Liang , Zhikang Niu , Wenxi Chen , Xie Chen

Target speech extraction (TSE) has achieved strong performance in relatively simple conditions such as one-speaker-plus-noise and two-speaker mixtures, but its performance remains unsatisfactory in noisy multi-speaker scenarios. To address…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-16 Ziling Huang , Junnan Wu , Lichun Fan , Zhenbo Luo , Jian Luan , Haixin Guan , Yanhua Long

Real-world audio recordings often contain multiple speakers and various degradations, which limit both the quantity and quality of speech data available for building state-of-the-art speech processing models. Although end-to-end approaches…

Sound · Computer Science 2026-01-27 Kohei Asai , Wataru Nakata , Yuki Saito , Hiroshi Saruwatari

MeanFlow (MF) is a diffusion-motivated generative model that enables efficient few-step generation by learning long jumps directly from noise to data. In practice, it is often used as a latent MF by leveraging the pre-trained Stable…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Zheyuan Hu , Chieh-Hsin Lai , Ge Wu , Yuki Mitsufuji , Stefano Ermon

Achieving robust and personalized performance in neuro-steered Target Speaker Extraction (TSE) remains a significant challenge for next-generation hearing aids. This is primarily due to two factors: the inherent non-stationarity of EEG…

Sound · Computer Science 2025-09-23 Qiushi Han , Yuan Liao , Youhao Si , Liya Huang

We introduce a one-step generative policy for offline reinforcement learning that maps noise directly to actions via a residual reformulation of MeanFlow, making it compatible with Q-learning. While one-step Gaussian policies enable fast…

Machine Learning · Computer Science 2025-11-18 Zeyuan Wang , Da Li , Yulin Chen , Ye Shi , Liang Bai , Tianyuan Yu , Yanwei Fu

The brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlook the issue of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-03 Cunhang Fan , Ying Chen , Jian Zhou , Zexu Pan , Jingjing Zhang , Youdian Gao , Xiaoke Yang , Zhengqi Wen , Zhao Lv

Language-queried target sound extraction (TSE) aims to extract specific sounds from mixtures based on language queries. Traditional fully-supervised training schemes require extensively annotated parallel audio-text data, which are…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-24 Hao Ma , Zhiyuan Peng , Xu Li , Yukai Li , Mingjie Shao , Qiuqiang Kong , Ju Liu

Speech enhancement (SE) based on diffusion probabilistic models has exhibited impressive performance, while requiring a relatively high number of function evaluations (NFE). Recently, SE based on flow matching has been proposed, which…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-20 Seonggyu Lee , Sein Cheong , Sangwook Han , Kihyuk Kim , Jong Won Shin

Diffusion models have shown strong performance in speech enhancement, but their real-time applicability has been limited by multi-step iterative sampling. Consistency distillation has recently emerged as a promising alternative by…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-19 Liang Xu , Longfei Felix Yan , W. Bastiaan Kleijn

We present the first neural target speech extraction (TSE) system that uses human feedback for iterative refinement. Our approach allows users to mark specific segments of the TSE output, generating an edit mask. The refinement system then…

Sound · Computer Science 2025-08-06 Malek Itani , Ashton Graves , Sefik Emre Eskimez , Shyamnath Gollakota

Audio-visual target speaker extraction (AV-TSE) models primarily rely on visual cues from the target speaker. However, humans also leverage linguistic knowledge, such as syntactic constraints, next word prediction, and prior knowledge of…

Sound · Computer Science 2025-11-11 Wenxuan Wu , Shuai Wang , Xixin Wu , Helen Meng , Haizhou Li

We propose DiffSep, a new single channel source separation method based on score-matching of a stochastic differential equation (SDE). We craft a tailored continuous time diffusion-mixing process starting from the separated sources and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-03 Robin Scheibler , Youna Ji , Soo-Whan Chung , Jaeuk Byun , Soyeon Choe , Min-Seok Choi

Target confusion, defined as occasional switching to non-target speakers, poses a key challenge for end-to-end speaker extraction (E2E-SE) systems. We argue that this problem is largely caused by the lack of generalizability and…

Sound · Computer Science 2025-05-29 Zhenghai You , Zhenyu Zhou , Lantian Li , Dong Wang

Few-step generation has been a long-standing goal, with recent one-step generation methods exemplified by MeanFlow achieving remarkable results. Existing research on MeanFlow primarily focuses on class-to-image generation. However, an…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Chenxi Zhao , Chen Zhu , Xiaokun Feng , Aiming Hao , Jiashu Zhu , Jiachen Lei , Jiahong Wu , Xiangxiang Chu , Jufeng Yang

The primary goal of multi-modal TSE (MTSE) is to extract a target speaker from a speech mixture using complementary information from different modalities, such as audio enrolment and visual feeds corresponding to the target speaker. MTSE…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-10 Srikanth Korse , Mohamed Elminshawi , Emanuel A. P. Habets , Srikanth Raj Chetupalli

Mean flow (MeanFlow) enables efficient, high-fidelity image generation, yet its single-function evaluation (1-NFE) generation often cannot yield compelling results. We address this issue by introducing RMFlow, an efficient multimodal…

Machine Learning · Computer Science 2026-02-03 Yuhao Huang , Shih-Hsin Wang , Andrea L. Bertozzi , Bao Wang

This paper investigates the use of relative cues for text-based target speech extraction (TSE). We first provide a theoretical justification for relative cues from the perspectives of human perception and label quantization, showing that…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-03 Wang Dai , Archontis Politis , Tuomas Virtanen
‹ Prev 1 3 4 5 6 7 10 Next ›