中文
相关论文

相关论文: Phase reconstruction from amplitude spectrograms b…

200 篇论文

This paper presents a novel phase reconstruction method (only from a given amplitude spectrogram) by combining a signal-processing-based approach and a deep neural network (DNN). To retrieve a time-domain signal from its amplitude…

声音 · 计算机科学 2019-03-12 Yoshiki Masuyama , Kohei Yatabe , Yuma Koizumi , Yasuhiro Oikawa , Noboru Harada

Phase reconstruction, which estimates phase from a given amplitude spectrogram, is an active research field in acoustical signal processing with many applications including audio synthesis. To take advantage of rich knowledge from data,…

音频与语音处理 · 电气工程与系统科学 2020-02-17 Yoshiki Masuyama , Kohei Yatabe , Yuma Koizumi , Yasuhiro Oikawa , Noboru Harada

In this paper, we address the problem of reconstructing a time-domain signal (or a phase spectrogram) solely from a magnitude spectrogram. Since magnitude spectrograms do not contain phase information, we must restore or infer phase…

信号处理 · 电气工程与系统科学 2018-04-09 Keisuke Oyamada , Hirokazu Kameoka , Takuhiro Kaneko , Kou Tanaka , Nobukatsu Hojo , Hiroyasu Ando

This paper proposes an approach to the joint modeling of the short-time Fourier transform magnitude and phase spectrograms with a deep generative model. We assume that the magnitude follows a Gaussian distribution and the phase follows a…

声音 · 计算机科学 2022-07-18 Aditya Arie Nugraha , Kouhei Sekiguchi , Kazuyoshi Yoshii

We propose an optimization-based method for reconstructing a time-domain signal from a low-dimensional spectral representation such as a mel-spectrogram. Phase reconstruction has been studied to reconstruct a time-domain signal from the…

声音 · 计算机科学 2023-07-25 Yoshiki Masuyama , Natsuki Ueno , Nobutaka Ono

Phase retrieval is a problem encountered not only in speech and audio processing, but in many other fields such as optics. Iterative algorithms based on non-convex set projections are effective and frequently used for retrieving the phase…

音频与语音处理 · 电气工程与系统科学 2022-11-10 Tal Peer , Simon Welker , Timo Gerkmann

This paper presents a novel speech phase prediction model which predicts wrapped phase spectra directly from amplitude spectra by neural networks. The proposed model is a cascade of a residual convolutional network and a parallel estimation…

声音 · 计算机科学 2023-02-17 Yang Ai , Zhen-Hua Ling

We propose a method for estimating the magnitude distribution of an acoustic field from spatially sparse magnitude measurements. Such a method is useful when phase measurements are unreliable or inaccessible. Physics-informed neural…

声音 · 计算机科学 2026-01-28 Karl Schrader , Shoichi Koyama , Tomohiko Nakamura , Mirco Pezzoli

This paper introduces a novel technique for reconstructing the phase of modified spectrograms of audio signals. From the analysis of mixtures of sinusoids we obtain relationships between phases of successive time frames in the…

声音 · 计算机科学 2016-05-25 Paul Magron , Roland Badeau , Bertrand David

We propose a neural network-based speech enhancement (SE) method called the phase-aware recurrent two stage network (rTSN). The rTSN is an extension of our previously proposed two stage network (TSN) framework. This TSN framework was…

音频与语音处理 · 电气工程与系统科学 2020-01-28 Juntae Kim , Jaesung Bae

In recent years generative adversarial network (GAN) based models have been successfully applied for unsupervised speech-to-speech conversion.The rich compact harmonic view of the magnitude spectrogram is considered a suitable choice for…

音频与语音处理 · 电气工程与系统科学 2020-05-20 Mohammad Asif Khan , Fabien Cardinaux , Stefan Uhlich , Marc Ferras , Asja Fischer

This paper presents a two-stage online phase reconstruction framework using causal deep neural networks (DNNs). Phase reconstruction is a task of recovering phase of the short-time Fourier transform (STFT) coefficients only from the…

声音 · 计算机科学 2022-11-16 Yoshiki Masuyama , Kohei Yatabe , Kento Nagatomo , Yasuhiro Oikawa

We present a Machine Learning-based method for tomographic reconstruction of dense layered objects, with range of projection angles limited to $\pm $10$^\circ$. Whereas previous approaches to phase tomography generally require two steps,…

图像与视频处理 · 电气工程与系统科学 2020-01-08 Alexandre Goy , Girish Rughoobur , Shuai Li , Kwabena Arthur , Akintunde I. Akinwande , George Barbastathis

We propose the multi-head convolutional neural network (MCNN) architecture for waveform synthesis from spectrograms. Nonlinear interpolation in MCNN is employed with transposed convolution layers in parallel heads. MCNN achieves more than…

声音 · 计算机科学 2018-12-26 Sercan O. Arik , Heewoo Jun , Gregory Diamos

Diffusion models are receiving a growing interest for a variety of signal generation tasks such as speech or music synthesis. WaveGrad, for example, is a successful diffusion model that conditionally uses the mel spectrogram to guide a…

声音 · 计算机科学 2024-02-27 Haocheng Liu , Teysir Baoueb , Mathieu Fontaine , Jonathan Le Roux , Gael Richard

Finite Rate of Innovation (FRI) sampling theory enables reconstruction of classes of continuous non-bandlimited signals that have a small number of free parameters from their low-rate discrete samples. This task is often translated into a…

信号处理 · 电气工程与系统科学 2023-07-21 Vincent C. H. Leung , Jun-Jie Huang , Yonina C. Eldar , Pier Luigi Dragotti

This paper proposes a deep neural network (DNN)-based multi-channel speech enhancement system in which a DNN is trained to maximize the quality of the enhanced time-domain signal. DNN-based multi-channel speech enhancement is often…

音频与语音处理 · 电气工程与系统科学 2020-02-17 Yoshiki Masuyama , Masahito Togami , Tatsuya Komatsu

Phase recovery of modified spectrograms is a major issue in audio signal processing applications, such as source separation. This paper introduces a novel technique for estimating the phases of components in complex mixtures within onset…

声音 · 计算机科学 2016-11-17 Paul Magron , Roland Badeau , Bertrand David

The Fourier transform, an explicit decomposition method for visual signals, has been employed to explain the out-of-distribution generalization behaviors of Deep Neural Networks (DNNs). Previous studies indicate that the amplitude spectrum…

计算机视觉与模式识别 · 计算机科学 2024-08-29 Chengming Hu , Yeqian Du , Rui Wang , Hao Chen , Congcong Zhu

This study investigates phase reconstruction for deep learning based monaural talker-independent speaker separation in the short-time Fourier transform (STFT) domain. The key observation is that, for a mixture of two sources, with their…

声音 · 计算机科学 2018-11-26 Zhong-Qiu Wang , Ke Tan , DeLiang Wang
‹ 上一页 1 2 3 10 下一页 ›