English
Related papers

Related papers: Independent Deeply Learned Matrix Analysis for Mul…

200 papers

Speech deepfake source verification systems aims to determine whether two synthetic speech utterances originate from the same source generator, often assuming that the resulting source embeddings are independent of speaker traits. However,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-24 Xi Xuan , Wenxin Zhang , Zhiyu Li , Jennifer Williams , Ville Hautamäki , Tomi H. Kinnunen

Learning how to localize and separate individual object sounds in the audio channel of the video is a difficult task. Current state-of-the-art methods predict audio masks from artificially mixed spectrograms, known as Mix-and-Separate…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Tanzila Rahman , Leonid Sigal

In recent years, music source separation has been one of the most intensively studied research areas in music information retrieval. Improvements in deep learning lead to a big progress in music source separation performance. However, most…

Sound · Computer Science 2019-08-20 Jie Hwan Lee , Hyeong-Seok Choi , Kyogu Lee

Non-negative Matrix Factorization (NMF) has already been applied to learn speaker characterizations from single or non-simultaneous speech for speaker recognition applications. It is also known for its good performance in (blind) source…

Sound · Computer Science 2016-05-02 Jeroen Zegers , Hugo Van hamme

A common challenge in the natural sciences is to disentangle distinct, unknown sources from observations. Examples of this source separation task include deblending galaxies in a crowded field, distinguishing the activity of individual…

Machine Learning · Computer Science 2025-10-08 Sebastian Wagner-Carena , Aizhan Akhmetzhanova , Sydney Erickson

Linear Independent Component Analysis (ICA) is a blind source separation technique that has been used in various domains to identify independent latent sources from observed signals. In order to obtain a higher signal-to-noise ratio, the…

Machine Learning · Computer Science 2023-12-04 Ambroise Heurtebise , Pierre Ablin , Alexandre Gramfort

We introduce a new information maximization (infomax) approach for the blind source separation problem. The proposed framework provides an information-theoretic perspective for determinant maximization-based structured matrix factorization…

Information Theory · Computer Science 2022-05-03 Alper T. Erdogan

We introduce PodcastMix, a dataset formalizing the task of separating background music and foreground speech in podcasts. We aim at defining a benchmark suitable for training and evaluating (deep learning) source separation models. To that…

Sound · Computer Science 2022-07-18 Nicolás Schmidt , Jordi Pons , Marius Miron

Recently, many methods based on deep learning have been proposed for music source separation. Some state-of-the-art methods have shown that stacking many layers with many skip connections improve the SDR performance. Although such a deep…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-25 Minseok Kim , Woosung Choi , Jaehwa Chung , Daewon Lee , Soonyoung Jung

In this paper we propose a method for separation of moving sound sources. The method is based on first tracking the sources and then estimation of source spectrograms using multichannel non-negative matrix factorization (NMF) and extracting…

Sound · Computer Science 2017-10-30 Joonas Nikunen , Aleksandr Diment , Tuomas Virtanen

We develop an end-to-end system for multi-channel, multi-speaker automatic speech recognition. We propose a frontend for joint source separation and dereverberation based on the independent vector analysis (IVA) paradigm. It uses the fast…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-04 Robin Scheibler , Wangyou Zhang , Xuankai Chang , Shinji Watanabe , Yanmin Qian

A novel extension of Independent Component and Independent Vector Analysis for blind extraction/separation of one or several sources from time-varying mixtures is proposed. The mixtures are assumed to be separable source-by-source in series…

Signal Processing · Electrical Eng. & Systems 2021-05-12 Zbyněk Koldovský , Václav Kautský , Petr Tichavský

This paper studies the density priors for independent vector analysis (IVA) with convolutive speech mixture separation as the exemplary application. Most existing source priors for IVA are too simplified to capture the fine structures of…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-07 Xi-Lin Li

Deep neural network based methods have been successfully applied to music source separation. They typically learn a mapping from a mixture spectrogram to a set of source spectrograms, all with magnitudes only. This approach has several…

Sound · Computer Science 2021-09-14 Qiuqiang Kong , Yin Cao , Haohe Liu , Keunwoo Choi , Yuxuan Wang

Perceiving a scene most fully requires all the senses. Yet modeling how objects look and sound is challenging: most natural scenes and events contain multiple objects, and the audio track mixes all the sound sources together. We propose to…

Computer Vision and Pattern Recognition · Computer Science 2018-07-27 Ruohan Gao , Rogerio Feris , Kristen Grauman

This study investigates phase reconstruction for deep learning based monaural talker-independent speaker separation in the short-time Fourier transform (STFT) domain. The key observation is that, for a mixture of two sources, with their…

Sound · Computer Science 2018-11-26 Zhong-Qiu Wang , Ke Tan , DeLiang Wang

We propose a knowledge-driven, model-based approach to segmenting audio into single-category and mixed-category chunks with applications to source separation. "Knowledge" here denotes information associated with the data, such as music…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-26 Chun-wei Ho , Sabato Marco Siniscalchi , Kai Li , Chin-Hui Lee

Online blind source separation is essential for both speech communication and human-machine interaction. Among existing approaches, overdetermined independent vector analysis (OverIVA) delivers strong performance by exploiting the…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Kang Chen , Xianrui Wang , Yichen Yang , Andreas Brendel , Gongping Huang , Zbyněk Koldovský , Jingdong Chen , Jacob Benesty , Shoji Makino

Language-queried Audio Source Separation (LASS) enables open-vocabulary sound separation via natural language queries. While existing methods rely on task-specific training, we explore whether pretrained diffusion models, originally…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-26 Geonyoung Lee , Geonhee Han , Paul Hongsuck Seo

We present a monophonic source separation system that is trained by only observing mixtures with no ground truth separation information. We use a deep clustering approach which trains on multi-channel mixtures and learns to project…

Machine Learning · Computer Science 2021-05-14 Efthymios Tzinis , Shrikant Venkataramani , Paris Smaragdis