中文
相关论文

相关论文: Mel-Spectrogram Inversion via Alternating Directio…

200 篇论文

A non-iterative method for the construction of the Short-Time Fourier Transform (STFT) phase from the magnitude is presented. The method is based on the direct relationship between the partial derivatives of the phase and the logarithm of…

声音 · 计算机科学 2019-03-27 Zdeněk Průša , Peter Balazs , Peter L. Søndergaard

Full waveform inversion (FWI) is beginning to be used to characterize weak seismic events at different scales, an example of which is microseismic event (MSE) characterization. However, FWI with unknown sources is a severely underdetermined…

最优化与控制 · 数学 2021-08-17 Hossein S. Aghamiry , Ali Gholami , Stéphane Operto , Alison Malcolm

This study proposes novel signal analysis methods for replay speech detection in automatic speaker verification (ASV) systems. The proposed methods -- arbitrary analysis (AA), mel scale analysis (MA), and constant Q analysis (CQA) -- are…

音频与语音处理 · 电气工程与系统科学 2025-05-28 Lee Shih Kuang

This paper introduces a novel technique for reconstructing the phase of modified spectrograms of audio signals. From the analysis of mixtures of sinusoids we obtain relationships between phases of successive time frames in the…

声音 · 计算机科学 2016-05-25 Paul Magron , Roland Badeau , Bertrand David

Multi-segment reconstruction (MSR) problem consists of recovering a signal from noisy segments with unknown positions of the observation windows. One example arises in DNA sequence assembly, which is typically solved by matching short reads…

信号处理 · 电气工程与系统科学 2018-02-27 Mona Zehni , Minh N. Do , Zhizhen Zhao

The alternating direction method of multipliers (ADMM) is one of the most widely used first-order optimisation methods in the literature owing to its simplicity, flexibility and efficiency. Over the years, numerous efforts are made to…

最优化与控制 · 数学 2019-12-02 Clarice Poon , Jingwei Liang

This paper proposes a novel neural denoising vocoder that can generate clean speech waveforms from noisy mel-spectrograms. The proposed neural denoising vocoder consists of two components, i.e., a spectrum predictor and a enhancement…

音频与语音处理 · 电气工程与系统科学 2024-11-20 Hui-Peng Du , Ye-Xin Lu , Yang Ai , Zhen-Hua Ling

The Alternating Direction Method of Multipliers (ADMM) provides a natural way of solving inverse problems with multiple partial differential equations (PDE) forward models and nonsmooth regularization. ADMM allows splitting these…

数值分析 · 数学 2021-04-29 Luke Lozenski , Umberto Villa

With active research in audio compression techniques yielding substantial breakthroughs, spectral reconstruction of low-quality audio waves remains a less indulged topic. In this paper, we propose a novel approach for reconstructing higher…

声音 · 计算机科学 2021-08-10 Darshan Deshpande , Harshavardhan Abichandani

On the heels of orthogonal time frequency space (OTFS) modulation, the recently discovered affine frequency division multiplexing (AFDM) is a promising waveform for the sixth-generation wireless network. In this paper, we study the…

信息论 · 计算机科学 2023-08-29 Haoran Yin , Xizhang Wei , Yanqun Tang , Kai Yang

While log-amplitude mel-spectrogram has widely been used as the feature representation for processing speech based on deep learning, the effectiveness of another aspect of speech spectrum, i.e., phase information, was shown recently for…

声音 · 计算机科学 2022-05-02 Shunsuke Hidaka , Kohei Wakamiya , Tokihiko Kaburagi

We propose an alternating direction method of multipliers (ADMM) to solve an optimization problem stemming from inverse lithography. The objective functional of the optimization problem includes three terms: the misfit between the imaging…

数值分析 · 数学 2026-04-21 Junqing Chen , Haibo Liu

This paper proposes a new loss using short-time Fourier transform (STFT) spectra for the aim of training a high-performance neural speech waveform model that predicts raw continuous speech waveform samples directly. Not only amplitude…

音频与语音处理 · 电气工程与系统科学 2018-10-31 Shinji Takaki , Toru Nakashika , Xin Wang , Junichi Yamagishi

In this paper, we focus on the approximation of smooth functions $f: [-\pi, \pi] \rightarrow \mathbb{C}$, up to an unresolvable global phase ambiguity, from a finite set of Short Time Fourier Transform (STFT) magnitude (i.e., spectrogram)…

数值分析 · 数学 2021-06-07 Mark Iwen , Michael Perlmutter , Nada Sissouno , Aditya Viswanathan

Deep neural networks have been applied to audio spectrograms for respiratory sound classification. Existing models often treat the spectrogram as a synthetic image while overlooking its physical characteristics. In this paper, a Multi-View…

声音 · 计算机科学 2024-05-31 Wentao He , Yuchen Yan , Jianfeng Ren , Ruibin Bai , Xudong Jiang

We propose a method for automatic local time-adaptation of the spectrogram of audio signals: it is based on the decomposition of a signal within a Gabor multi-frame through the STFT operator. The sparsity of the analysis in every individual…

声音 · 计算机科学 2011-09-29 M. Liuni , A. Röbel , M. Romito , X. Rodet

Previous speech enhancement methods focus on estimating the short-time spectrum of speech signals due to its short-term stability. However, these methods often only estimate the clean magnitude spectrum and reuse the noisy phase when…

声音 · 计算机科学 2019-10-23 Chuang Geng , Lei Wang

This paper introduces an innovative approach for signal reconstruction using data acquired through multi-input-multi-output (MIMO) sampling. First, we show that it is possible to perfectly reconstruct a set of periodic band-limited signals…

信息论 · 计算机科学 2024-09-09 Dong Cheng , Xiaoxiao Hu , Kit Ian Kou

Time-frequency representation (TFR) allowing for mode reconstruction plays a significant role in interpreting and analyzing the nonstationary signal constituted of various modes. However, it is difficult for most previous methods to handle…

信号处理 · 电气工程与系统科学 2021-09-01 Haijian Zhang , Guang Hua

In this work, we propose CleanMel, a single-channel Mel-spectrogram denoising and dereverberation network for improving both speech quality and automatic speech recognition (ASR) performance. The proposed network takes as input the noisy…

音频与语音处理 · 电气工程与系统科学 2025-07-31 Nian Shao , Rui Zhou , Pengyu Wang , Xian Li , Ying Fang , Yujie Yang , Xiaofei Li