English
Related papers

Related papers: Model-based STFT phase recovery for audio source s…

200 papers

Analytical methods are fundamental in studying acoustics problems. One of the important tools is the Wiener-Hopf method, which can be used to solve many canonical problems with sharp transitions in boundary conditions on a plane/plate.…

Numerical Analysis · Mathematics 2024-03-27 Matthew Nethercote , Anastasia Kisil , Raphael Assier

We introduce the Latent Fourier Transform (LatentFT), a framework that provides novel frequency-domain controls for generative music models. LatentFT combines a diffusion autoencoder with a latent-space Fourier transform to separate musical…

Sound · Computer Science 2026-04-21 Mason Wang , Cheng-Zhi Anna Huang

In this paper, we present an assortment of both standard and advanced Fourier techniques that are useful in the analysis of astrophysical time series of very long duration -- where the observation time is much greater than the time…

Astrophysics · Physics 2009-11-07 Scott M. Ransom , Stephen S. Eikenberry , John Middleditch

We consider the recovery of a continuous-time Wiener process from a quantized or lossy compressed version of its uniform samples under limited bitrate and sampling rate. We derive a closed form expression for the optimal tradeoff among…

Information Theory · Computer Science 2018-07-27 Alon Kipnis , Andrea J. Goldsmith , Yonina C. Eldar

This paper addresses the problem of separating audio sources from time-varying convolutive mixtures. We propose a probabilistic framework based on the local complex-Gaussian model combined with non-negative matrix factorization. The…

Deep neural network based methods have been successfully applied to music source separation. They typically learn a mapping from a mixture spectrogram to a set of source spectrograms, all with magnitudes only. This approach has several…

Sound · Computer Science 2021-09-14 Qiuqiang Kong , Yin Cao , Haohe Liu , Keunwoo Choi , Yuxuan Wang

In this paper, we derive a new class of methods for the classic 2D phase unwrapping problem of recovering a phase function from its wrapped form. For this, we consider the wrapped phase as a wavefront aberration in an optical system, and…

Numerical Analysis · Mathematics 2025-03-14 Simon Hubmer , Victoria Laidlaw , Ronny Ramlau , Ekaterina Sherina , Bernadett Stadler

Deep learning-based techniques for automatic dysarthric speech detection have recently attracted interest in the research community. State-of-the-art techniques typically learn neurotypical and dysarthric discriminative representations by…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-04 Ina Kodrasi

This paper presents FastFit, a novel neural vocoder architecture that replaces the U-Net encoder with multiple short-time Fourier transforms (STFTs) to achieve faster generation rates without sacrificing sample quality. We replaced each…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-19 Won Jang , Dan Lim , Heayoung Park

In this paper, we proposed a robust music genre classification method based on a sparse FFT based feature extraction method which extracted with discriminating power of spectral analysis of non-stationary audio signals, and the capability…

Sound · Computer Science 2018-03-14 Mehdi Banitalebi-Dehkordi , Amin Banitalebi-Dehkordi

Convolutional neural networks (CNN) are widely used for speech emotion recognition (SER). In such cases, the short time fourier transform (STFT) spectrogram is the most popular choice for representing speech, which is fed as input to the…

Audio and Speech Processing · Electrical Eng. & Systems 2019-08-09 Shruti Gupta , Md. Shah Fahad , Akshay Deepak

We propose a method using a long short-term memory (LSTM) network to estimate the noise power spectral density (PSD) of single-channel audio signals represented in the short time Fourier transform (STFT) domain. An LSTM network common to…

Signal Processing · Electrical Eng. & Systems 2020-11-11 Xiaofei Li , Simon Leglaive , Laurent Girin , Radu Horaud

We introduce an audio texture synthesis algorithm based on scattering moments. A scattering transform is computed by iteratively decomposing a signal with complex wavelet filter banks and computing their amplitude envelop. Scattering…

Applications · Statistics 2013-11-05 Joan Bruna , Stéphane Mallat

We present a simple, frequency domain, preprocessing step to Kirchhoff migration that allows the method to image scatterers when the wave field phase information is lost at the receivers, and only intensities are measured. The resulting…

Numerical Analysis · Mathematics 2016-09-21 Patrick Bardsley , Fernando Guevara Vasquez

This article introduces a new parametric synthesis method for sound textures based on existing works in visual and sound texture synthesis. Starting from a base sound signal, an optimization process is performed until the cross-correlations…

Sound · Computer Science 2019-10-22 Hugo Caracalla , Axel Roebel

Separating audio mixtures into individual instrument tracks has been a long standing challenging task. We introduce a novel weakly supervised audio source separation approach based on deep adversarial learning. Specifically, our loss…

Sound · Computer Science 2018-05-18 Ning Zhang , Junchi Yan , Yuchen Zhou

With the growing demand for non-Euclidean data analysis, graph signal processing (GSP) has gained significant attention for its capability to handle complex time-varying data. This paper introduces a novel sampling method based on the joint…

General Mathematics · Mathematics 2025-06-03 Yu Zhang , Bing-Zhao Li

Decomposition of an audio mixture into harmonic and percussive components, namely harmonic/percussive source separation (HPSS), is a useful pre-processing tool for many audio applications. Popular approaches to HPSS exploit the distinctive…

Audio and Speech Processing · Electrical Eng. & Systems 2019-03-14 Yoshiki Masuyama , Kohei Yatabe , Yasuhiro Oikawa

Signal reconstruction from its mel-spectrogram is known as mel-spectrogram inversion and has many applications, including speech and foley sound synthesis. In this paper, we propose a mel-spectrogram inversion method based on a rigorous…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-14 Yoshiki Masuyama , Natsuki Ueno , Nobutaka Ono

Synchrosqueezing transform (SST) is a useful tool for vibration signal analysis due to its high time-frequency (TF) concentration and reconstruction properties. However, existing SST requires much processing time for large-scale data. In…

Signal Processing · Electrical Eng. & Systems 2020-03-17 Dong He , Hongrui Cao