English
Related papers

Related papers: Phase reconstruction from amplitude spectrograms b…

200 papers

This paper presents a novel phase reconstruction method (only from a given amplitude spectrogram) by combining a signal-processing-based approach and a deep neural network (DNN). To retrieve a time-domain signal from its amplitude…

Sound · Computer Science 2019-03-12 Yoshiki Masuyama , Kohei Yatabe , Yuma Koizumi , Yasuhiro Oikawa , Noboru Harada

Phase reconstruction, which estimates phase from a given amplitude spectrogram, is an active research field in acoustical signal processing with many applications including audio synthesis. To take advantage of rich knowledge from data,…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-17 Yoshiki Masuyama , Kohei Yatabe , Yuma Koizumi , Yasuhiro Oikawa , Noboru Harada

In this paper, we address the problem of reconstructing a time-domain signal (or a phase spectrogram) solely from a magnitude spectrogram. Since magnitude spectrograms do not contain phase information, we must restore or infer phase…

Signal Processing · Electrical Eng. & Systems 2018-04-09 Keisuke Oyamada , Hirokazu Kameoka , Takuhiro Kaneko , Kou Tanaka , Nobukatsu Hojo , Hiroyasu Ando

This paper proposes an approach to the joint modeling of the short-time Fourier transform magnitude and phase spectrograms with a deep generative model. We assume that the magnitude follows a Gaussian distribution and the phase follows a…

Sound · Computer Science 2022-07-18 Aditya Arie Nugraha , Kouhei Sekiguchi , Kazuyoshi Yoshii

We propose an optimization-based method for reconstructing a time-domain signal from a low-dimensional spectral representation such as a mel-spectrogram. Phase reconstruction has been studied to reconstruct a time-domain signal from the…

Sound · Computer Science 2023-07-25 Yoshiki Masuyama , Natsuki Ueno , Nobutaka Ono

Phase retrieval is a problem encountered not only in speech and audio processing, but in many other fields such as optics. Iterative algorithms based on non-convex set projections are effective and frequently used for retrieving the phase…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-10 Tal Peer , Simon Welker , Timo Gerkmann

This paper presents a novel speech phase prediction model which predicts wrapped phase spectra directly from amplitude spectra by neural networks. The proposed model is a cascade of a residual convolutional network and a parallel estimation…

Sound · Computer Science 2023-02-17 Yang Ai , Zhen-Hua Ling

We propose a method for estimating the magnitude distribution of an acoustic field from spatially sparse magnitude measurements. Such a method is useful when phase measurements are unreliable or inaccessible. Physics-informed neural…

Sound · Computer Science 2026-01-28 Karl Schrader , Shoichi Koyama , Tomohiko Nakamura , Mirco Pezzoli

This paper introduces a novel technique for reconstructing the phase of modified spectrograms of audio signals. From the analysis of mixtures of sinusoids we obtain relationships between phases of successive time frames in the…

Sound · Computer Science 2016-05-25 Paul Magron , Roland Badeau , Bertrand David

We propose a neural network-based speech enhancement (SE) method called the phase-aware recurrent two stage network (rTSN). The rTSN is an extension of our previously proposed two stage network (TSN) framework. This TSN framework was…

Audio and Speech Processing · Electrical Eng. & Systems 2020-01-28 Juntae Kim , Jaesung Bae

In recent years generative adversarial network (GAN) based models have been successfully applied for unsupervised speech-to-speech conversion.The rich compact harmonic view of the magnitude spectrogram is considered a suitable choice for…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-20 Mohammad Asif Khan , Fabien Cardinaux , Stefan Uhlich , Marc Ferras , Asja Fischer

This paper presents a two-stage online phase reconstruction framework using causal deep neural networks (DNNs). Phase reconstruction is a task of recovering phase of the short-time Fourier transform (STFT) coefficients only from the…

Sound · Computer Science 2022-11-16 Yoshiki Masuyama , Kohei Yatabe , Kento Nagatomo , Yasuhiro Oikawa

We present a Machine Learning-based method for tomographic reconstruction of dense layered objects, with range of projection angles limited to $\pm $10$^\circ$. Whereas previous approaches to phase tomography generally require two steps,…

Image and Video Processing · Electrical Eng. & Systems 2020-01-08 Alexandre Goy , Girish Rughoobur , Shuai Li , Kwabena Arthur , Akintunde I. Akinwande , George Barbastathis

We propose the multi-head convolutional neural network (MCNN) architecture for waveform synthesis from spectrograms. Nonlinear interpolation in MCNN is employed with transposed convolution layers in parallel heads. MCNN achieves more than…

Sound · Computer Science 2018-12-26 Sercan O. Arik , Heewoo Jun , Gregory Diamos

Diffusion models are receiving a growing interest for a variety of signal generation tasks such as speech or music synthesis. WaveGrad, for example, is a successful diffusion model that conditionally uses the mel spectrogram to guide a…

Sound · Computer Science 2024-02-27 Haocheng Liu , Teysir Baoueb , Mathieu Fontaine , Jonathan Le Roux , Gael Richard

Finite Rate of Innovation (FRI) sampling theory enables reconstruction of classes of continuous non-bandlimited signals that have a small number of free parameters from their low-rate discrete samples. This task is often translated into a…

Signal Processing · Electrical Eng. & Systems 2023-07-21 Vincent C. H. Leung , Jun-Jie Huang , Yonina C. Eldar , Pier Luigi Dragotti

This paper proposes a deep neural network (DNN)-based multi-channel speech enhancement system in which a DNN is trained to maximize the quality of the enhanced time-domain signal. DNN-based multi-channel speech enhancement is often…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-17 Yoshiki Masuyama , Masahito Togami , Tatsuya Komatsu

Phase recovery of modified spectrograms is a major issue in audio signal processing applications, such as source separation. This paper introduces a novel technique for estimating the phases of components in complex mixtures within onset…

Sound · Computer Science 2016-11-17 Paul Magron , Roland Badeau , Bertrand David

The Fourier transform, an explicit decomposition method for visual signals, has been employed to explain the out-of-distribution generalization behaviors of Deep Neural Networks (DNNs). Previous studies indicate that the amplitude spectrum…

Computer Vision and Pattern Recognition · Computer Science 2024-08-29 Chengming Hu , Yeqian Du , Rui Wang , Hao Chen , Congcong Zhu

This study investigates phase reconstruction for deep learning based monaural talker-independent speaker separation in the short-time Fourier transform (STFT) domain. The key observation is that, for a mixture of two sources, with their…

Sound · Computer Science 2018-11-26 Zhong-Qiu Wang , Ke Tan , DeLiang Wang
‹ Prev 1 2 3 10 Next ›