English
Related papers

Related papers: Mel-Spectrogram Inversion via Alternating Directio…

200 papers

A non-iterative method for the construction of the Short-Time Fourier Transform (STFT) phase from the magnitude is presented. The method is based on the direct relationship between the partial derivatives of the phase and the logarithm of…

Sound · Computer Science 2019-03-27 Zdeněk Průša , Peter Balazs , Peter L. Søndergaard

Full waveform inversion (FWI) is beginning to be used to characterize weak seismic events at different scales, an example of which is microseismic event (MSE) characterization. However, FWI with unknown sources is a severely underdetermined…

Optimization and Control · Mathematics 2021-08-17 Hossein S. Aghamiry , Ali Gholami , Stéphane Operto , Alison Malcolm

This study proposes novel signal analysis methods for replay speech detection in automatic speaker verification (ASV) systems. The proposed methods -- arbitrary analysis (AA), mel scale analysis (MA), and constant Q analysis (CQA) -- are…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-28 Lee Shih Kuang

This paper introduces a novel technique for reconstructing the phase of modified spectrograms of audio signals. From the analysis of mixtures of sinusoids we obtain relationships between phases of successive time frames in the…

Sound · Computer Science 2016-05-25 Paul Magron , Roland Badeau , Bertrand David

Multi-segment reconstruction (MSR) problem consists of recovering a signal from noisy segments with unknown positions of the observation windows. One example arises in DNA sequence assembly, which is typically solved by matching short reads…

Signal Processing · Electrical Eng. & Systems 2018-02-27 Mona Zehni , Minh N. Do , Zhizhen Zhao

The alternating direction method of multipliers (ADMM) is one of the most widely used first-order optimisation methods in the literature owing to its simplicity, flexibility and efficiency. Over the years, numerous efforts are made to…

Optimization and Control · Mathematics 2019-12-02 Clarice Poon , Jingwei Liang

This paper proposes a novel neural denoising vocoder that can generate clean speech waveforms from noisy mel-spectrograms. The proposed neural denoising vocoder consists of two components, i.e., a spectrum predictor and a enhancement…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-20 Hui-Peng Du , Ye-Xin Lu , Yang Ai , Zhen-Hua Ling

The Alternating Direction Method of Multipliers (ADMM) provides a natural way of solving inverse problems with multiple partial differential equations (PDE) forward models and nonsmooth regularization. ADMM allows splitting these…

Numerical Analysis · Mathematics 2021-04-29 Luke Lozenski , Umberto Villa

With active research in audio compression techniques yielding substantial breakthroughs, spectral reconstruction of low-quality audio waves remains a less indulged topic. In this paper, we propose a novel approach for reconstructing higher…

Sound · Computer Science 2021-08-10 Darshan Deshpande , Harshavardhan Abichandani

On the heels of orthogonal time frequency space (OTFS) modulation, the recently discovered affine frequency division multiplexing (AFDM) is a promising waveform for the sixth-generation wireless network. In this paper, we study the…

Information Theory · Computer Science 2023-08-29 Haoran Yin , Xizhang Wei , Yanqun Tang , Kai Yang

While log-amplitude mel-spectrogram has widely been used as the feature representation for processing speech based on deep learning, the effectiveness of another aspect of speech spectrum, i.e., phase information, was shown recently for…

Sound · Computer Science 2022-05-02 Shunsuke Hidaka , Kohei Wakamiya , Tokihiko Kaburagi

We propose an alternating direction method of multipliers (ADMM) to solve an optimization problem stemming from inverse lithography. The objective functional of the optimization problem includes three terms: the misfit between the imaging…

Numerical Analysis · Mathematics 2026-04-21 Junqing Chen , Haibo Liu

This paper proposes a new loss using short-time Fourier transform (STFT) spectra for the aim of training a high-performance neural speech waveform model that predicts raw continuous speech waveform samples directly. Not only amplitude…

Audio and Speech Processing · Electrical Eng. & Systems 2018-10-31 Shinji Takaki , Toru Nakashika , Xin Wang , Junichi Yamagishi

In this paper, we focus on the approximation of smooth functions $f: [-\pi, \pi] \rightarrow \mathbb{C}$, up to an unresolvable global phase ambiguity, from a finite set of Short Time Fourier Transform (STFT) magnitude (i.e., spectrogram)…

Numerical Analysis · Mathematics 2021-06-07 Mark Iwen , Michael Perlmutter , Nada Sissouno , Aditya Viswanathan

Deep neural networks have been applied to audio spectrograms for respiratory sound classification. Existing models often treat the spectrogram as a synthetic image while overlooking its physical characteristics. In this paper, a Multi-View…

Sound · Computer Science 2024-05-31 Wentao He , Yuchen Yan , Jianfeng Ren , Ruibin Bai , Xudong Jiang

We propose a method for automatic local time-adaptation of the spectrogram of audio signals: it is based on the decomposition of a signal within a Gabor multi-frame through the STFT operator. The sparsity of the analysis in every individual…

Sound · Computer Science 2011-09-29 M. Liuni , A. Röbel , M. Romito , X. Rodet

Previous speech enhancement methods focus on estimating the short-time spectrum of speech signals due to its short-term stability. However, these methods often only estimate the clean magnitude spectrum and reuse the noisy phase when…

Sound · Computer Science 2019-10-23 Chuang Geng , Lei Wang

This paper introduces an innovative approach for signal reconstruction using data acquired through multi-input-multi-output (MIMO) sampling. First, we show that it is possible to perfectly reconstruct a set of periodic band-limited signals…

Information Theory · Computer Science 2024-09-09 Dong Cheng , Xiaoxiao Hu , Kit Ian Kou

Time-frequency representation (TFR) allowing for mode reconstruction plays a significant role in interpreting and analyzing the nonstationary signal constituted of various modes. However, it is difficult for most previous methods to handle…

Signal Processing · Electrical Eng. & Systems 2021-09-01 Haijian Zhang , Guang Hua

In this work, we propose CleanMel, a single-channel Mel-spectrogram denoising and dereverberation network for improving both speech quality and automatic speech recognition (ASR) performance. The proposed network takes as input the noisy…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-31 Nian Shao , Rui Zhou , Pengyu Wang , Xian Li , Ying Fang , Yujie Yang , Xiaofei Li