English
Related papers

Related papers: A Comparison and Combination of Unsupervised Blind…

200 papers

Voice activity detection (VAD) is an important pre-processing step for speech technology applications. The task consists of deriving segment boundaries of audio signals which contain voicing information. In recent years, it has been shown…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-28 Eklavya Sarkar , RaviShankar Prasad , Mathew Magimai. -Doss

Learning how objects sound from video is challenging, since they often heavily overlap in a single audio channel. Current methods for visually-guided audio source separation sidestep the issue by training with artificially mixed video…

Computer Vision and Pattern Recognition · Computer Science 2019-08-22 Ruohan Gao , Kristen Grauman

In the community of remote sensing, nonlinear mixing models have recently received particular attention in hyperspectral image processing. In this paper, we present a novel nonlinear spectral unmixing method following the recent multilinear…

Computer Vision and Pattern Recognition · Computer Science 2017-10-11 Qi Wei , Marcus Chen , Jean-Yves Tourneret , Simon Godsill

Significant challenges exist in efficient data analysis of most advanced experimental and observational techniques because the collected signals often include unwanted contributions--such as background and signal distortions--that can…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Yuan Ni , Zhantao Chen , Alexander N. Petsch , Edmund Xu , Cheng Peng , Alexander I. Kolesnikov , Sugata Chowdhury , Arun Bansil , Jana B. Thayer , Joshua J. Turner

In this work, we present a method for learning interpretable music signal representations directly from waveform signals. Our method can be trained using unsupervised objectives and relies on the denoising auto-encoder model that uses a…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-02 Stylianos I. Mimilakis , Konstantinos Drossos , Gerald Schuller

Unsupervised representation learning for speech processing has matured greatly in the last few years. Work in computer vision and natural language processing has paved the way, but speech data offers unique challenges. As a result, methods…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-04 Lasse Borgholt , Jakob Drachmann Havtorn , Joakim Edin , Lars Maaløe , Christian Igel

Modeling non Gaussian and non stationary signals and images has always been one of the most important part of signal and image processing methods. In this paper, first we propose a few new models, all based on using hidden variables for…

Data Analysis, Statistics and Probability · Physics 2007-05-23 Ali Mohammad-Djafari

In the mobile communication field, some of the video applications boosted the interest of robust methods for video quality assessment. Out of all existing methods, We Preferred, No Reference Video Quality Assessment is the one which is most…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Amitesh Kumar Singam , Benny Lövström , Wlodek J. Kulesza

Unsupervised video object segmentation has often been tackled by methods based on recurrent neural networks and optical flow. Despite their complexity, these kinds of approaches tend to favour short-term temporal dependencies and are thus…

Computer Vision and Pattern Recognition · Computer Science 2019-10-25 Zhao Yang , Qiang Wang , Luca Bertinetto , Weiming Hu , Song Bai , Philip H. S. Torr

The independent low-rank matrix analysis (ILRMA) method stands out as a prominent technique for multichannel blind audio source separation. It leverages nonnegative matrix factorization (NMF) and nonnegative canonical polyadic decomposition…

Sound · Computer Science 2024-05-07 Jianyu Wang , Shanzheng Guan

We present a monophonic source separation system that is trained by only observing mixtures with no ground truth separation information. We use a deep clustering approach which trains on multi-channel mixtures and learns to project…

Machine Learning · Computer Science 2021-05-14 Efthymios Tzinis , Shrikant Venkataramani , Paris Smaragdis

We present a joint audio-visual model for isolating a single speech signal from a mixture of sounds such as other speakers and background noise. Solving this task using only audio as input is extremely challenging and does not provide an…

We propose an unsupervised approach for training separation models from scratch using RemixIT and Self-Remixing, which are recently proposed self-supervised learning methods for refining pre-trained models. They first separate mixtures with…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-04 Kohei Saijo , Tetsuji Ogawa

Although deep learning based multi-channel speech enhancement has achieved significant advancements, its practical deployment is often limited by constrained computational resources, particularly in low signal-to-noise ratio (SNR)…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-27 Zheng Wang , Xiaobin Rong , Yu Sun , Tianchi Sun , Zhibin Lin , Jing Lu

Radio frequency sources are observed at a fusion center via sensor measurements made over slow flat-fading channels. The number of sources may be larger than the number of sensors, but their activity is sparse and intermittent with bursty…

Signal Processing · Electrical Eng. & Systems 2019-08-07 Annan Dong , Osvaldo Simeone , Alexander Haimovich , Jason Dabin

We propose SAHMM-VAE, a source-wise adaptive Hidden Markov prior variational autoencoder for unsupervised blind source separation. Instead of treating the latent prior as a single generic regularizer, the proposed framework assigns each…

Machine Learning · Statistics 2026-03-30 Yuan-Hao Wei

Unsupervised single-channel overlapped speech recognition is one of the hardest problems in automatic speech recognition (ASR). Permutation invariant training (PIT) is a state of the art model-based approach, which applies a single neural…

Computation and Language · Computer Science 2017-12-27 Zhehuai Chen , Jasha Droppo , Jinyu Li , Wayne Xiong

Multichannel blind audio source separation aims to recover the latent sources from their multichannel mixtures without supervised information. One state-of-the-art blind audio source separation method, named independent low-rank matrix…

Sound · Computer Science 2021-03-31 Jianyu Wang , Shanzheng Guan , Shupei Liu , Xiao-Lei Zhang

Hyperspectral unmixing is a blind source separation problem which consists in estimating the reference spectral signatures contained in a hyperspectral image, as well as their relative contribution to each pixel according to a given mixture…

Data Analysis, Statistics and Probability · Physics 2017-11-21 Pierre-Antoine Thouvenin , Nicolas Dobigeon , Jean-Yves Tourneret

We propose a new estimation method for the blind source separation model of Bachoc et al. (2020). The new estimation is based on an eigenanalysis of a positive definite matrix defined in terms of multiple normalized spatial local covariance…

Methodology · Statistics 2022-08-29 Bo Zhang , Sixing Hao , Qiwei Yao