English
Related papers

Related papers: MAPSS: Manifold-based Assessment of Perceptual Sou…

200 papers

Music inpainting aims to reconstruct missing segments of a corrupted recording. While diffusion-based generative models improve reconstruction for medium-length gaps, they often struggle to preserve musical plausibility over multi-second…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-23 Sean Turland , Eloi Moliner , Vesa Välimäki

The matrix pencil method (MPM) is a well-known technique for estimating the parameters of exponentially damped sinusoids in noise by solving a generalized eigenvalue problem. However, in several cases, this is an ill-conditioned problem…

Signal Processing · Electrical Eng. & Systems 2024-04-18 Raymundo Albert , Cecilia G. Galarza

Denoising Diffusion Probabilistic Models (DDPM) are powerful state-of-the-art methods used to generate synthetic data from high-dimensional data distributions and are widely used for image, audio, and video generation as well as many more…

Machine Learning · Statistics 2025-04-25 Iskander Azangulov , George Deligiannidis , Judith Rousseau

Incompatibility of image descriptor and ranking is always neglected in image retrieval. In this paper, manifold learning and Gestalt psychology theory are involved to solve the incompatibility problem. A new holistic descriptor called…

Computer Vision and Pattern Recognition · Computer Science 2016-09-27 Shenglan Liu , Jun Wu , Lin Feng , Yang Liu , Hong Qiao , Wenbo Luo Muxin Sun , Wei Wang

Separating an audio scene into isolated sources is a fundamental problem in computer audition, analogous to image segmentation in visual scene analysis. Source separation systems based on deep learning are currently the most successful…

Sound · Computer Science 2018-11-07 Prem Seetharaman , Gordon Wichern , Jonathan Le Roux , Bryan Pardo

Existing learning-based multi-view stereo (MVS) methods rely on the depth range to build the 3D cost volume and may fail when the range is too large or unreliable. To address this problem, we propose a disparity-based MVS method based on…

Computer Vision and Pattern Recognition · Computer Science 2022-12-06 Qingsong Yan , Qiang Wang , Kaiyong Zhao , Bo Li , Xiaowen Chu , Fei Deng

The problem of mixed signals occurs in many different contexts; one of the most familiar being acoustics. The forward problem in acoustics consists of finding the sound pressure levels at various detectors resulting from sound signals…

Data Analysis, Statistics and Probability · Physics 2007-05-23 Kevin H. Knuth

In this paper we propose an efficient deep learning encoder-decoder network for performing Harmonic-Percussive Source Separation (HPSS). It is shown that we are able to greatly reduce the number of model trainable parameters by using a…

Sound · Computer Science 2019-07-31 Carlos Lordelo , Emmanouil Benetos , Simon Dixon , Sven Ahlbäck

Learning mappings of data on manifolds is an important topic in contemporary machine learning, with applications in astrophysics, geophysics, statistical physics, medical diagnosis, biochemistry, 3D object analysis. This paper studies the…

Numerical Analysis · Mathematics 2020-07-21 Guido Montúfar , Yu Guang Wang

Due to the enormous requirement in public security and intelligent transportation system, searching an identical vehicle has become more and more important. Current studies usually treat vehicle as an integral object and then train a…

Computer Vision and Pattern Recognition · Computer Science 2019-11-12 Ya Sun , Minxian Li , Jianfeng Lu

This paper proposes APSS, a novel neural speech separation model with parallel amplitude and phase spectrum estimation. Unlike most existing speech separation methods, the APSS distinguishes itself by explicitly estimating the phase…

Sound · Computer Science 2025-09-18 Fei Liu , Yang Ai , Zhen-Hua Ling

We consider a collection of $n$ points in $\mathbb{R}^d$ measured at $m$ times, which are encoded in an $n \times d \times m$ data tensor. Our objective is to define a single embedding of the $n$ points into Euclidean space which summarizes…

Classical Analysis and ODEs · Mathematics 2019-11-27 Nicholas F. Marshall , Matthew J. Hirn

We introduce a framework for audio source separation using embeddings on a hyperbolic manifold that compactly represent the hierarchical relationship between sound sources and time-frequency features. Inspired by recent successes modeling…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-12 Darius Petermann , Gordon Wichern , Aswin Subramanian , Jonathan Le Roux

The objective of deep learning methods based on encoder-decoder architectures for music source separation is to approximate either ideal time-frequency masks or spectral representations of the target music source(s). The spectral…

Perspective distortion (PD) causes unprecedented changes in shape, size, orientation, angles, and other spatial relationships of visual concepts in images. Precisely estimating camera intrinsic and extrinsic parameters is a challenging task…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Prakash Chandra Chhipa , Meenakshi Subhash Chippa , Kanjar De , Rajkumar Saini , Marcus Liwicki , Mubarak Shah

The audio source separation tasks, such as speech enhancement, speech separation, and music source separation, have achieved impressive performance in recent studies. The powerful modeling capabilities of deep neural networks give us hope…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-15 Lu Zhang , Chenxing Li , Feng Deng , Xiaorui Wang

Point Source (PS) detection is an important issue for future Cosmic Microwave Background (CMB) experiments since they are one of the main contaminants to the recovery of CMB signal at small scales. Improving its multifrequency detection…

Instrumentation and Methods for Astrophysics · Physics 2022-02-09 J. M. Casas , J. González-Nuevo , L. Bonavera , D. Herranz , S. L. Suárez Gómez , M. M. Cueli , D. Crespo , J. D. Santos , M. L. Sánchez , F. Sánchez-Lasheras , F. J. de Cos

Source separation is one of the signal processing's main emerging domain. Many techniques such as maximum likelihood (ML), Infomax, cumulant matching, estimating function, etc. have been used to address this difficult problem.…

Mathematical Physics · Physics 2009-10-31 Ali Mohammad-Djafari

Current metrics for text-to-image models typically rely on statistical metrics which inadequately represent the real preference of humans. Although recent work attempts to learn these preferences via human annotated images, they reduce the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Sixian Zhang , Bohan Wang , Junqiang Wu , Yan Li , Tingting Gao , Di Zhang , Zhongyuan Wang

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Hyeonggon Ryu , Seongyu Kim , Joon Son Chung , Arda Senocak