English
Related papers

Related papers: Mel-Spectrogram Inversion via Alternating Directio…

200 papers

Flow matching offers a robust and stable approach to training diffusion models. However, directly applying flow matching to neural vocoders can result in subpar audio quality. In this work, we present WaveFM, a reparameterized flow matching…

Sound · Computer Science 2025-03-24 Tianze Luo , Xingchen Miao , Wenbo Duan

We propose a method using a long short-term memory (LSTM) network to estimate the noise power spectral density (PSD) of single-channel audio signals represented in the short time Fourier transform (STFT) domain. An LSTM network common to…

Signal Processing · Electrical Eng. & Systems 2020-11-11 Xiaofei Li , Simon Leglaive , Laurent Girin , Radu Horaud

We demonstrate machine learning (ML) based reconstruction of terahertz transmission spectra using an electrically tunable grating-gate AlGaN/GaN plasmonic-crystal analyzer. The analyzer encodes the transmission spectrum into a…

Optics · Physics 2026-05-06 A. Witkowska , M. Dub , P. Sai , P. Tiwari , M. Sakowicz , J. A. Majewski , W. Knap

Speaker separation refers to isolating speech of interest in a multi-talker environment. Most methods apply real-valued Time-Frequency (T-F) masks to the mixture Short-Time Fourier Transform (STFT) to reconstruct the clean speech. Hence…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-16 Zhaoheng Ni , Michael I Mandel

The Alternating Direction Method of Multipliers (ADMM) has been studied for years. The traditional ADMM algorithm needs to compute, at each iteration, an (empirical) expected loss function on all training examples, resulting in a…

Machine Learning · Statistics 2014-06-10 Peilin Zhao , Jinwei Yang , Tong Zhang , Ping Li

Sparse signal recovery based on nonconvex and nonsmooth optimization problems has significant applications and demonstrates superior performance in signal processing and machine learning. This work deals with a scale-invariant…

Optimization and Control · Mathematics 2025-09-29 Lang Yu , Nanjing Huang

We study a class of structured convex optimization problems, which have a two-block separable objective and nonlinear functional constraints as well as affine constraints that couple the two block variables. Such problems naturally arise…

Optimization and Control · Mathematics 2026-02-27 Zhengjie Xiong , Yangyang Xu

Decoding speech from brain signals is a challenging research problem. Although existing technologies have made progress in reconstructing the mel spectrograms of auditory stimuli at the word or letter level, there remain core challenges in…

Sound · Computer Science 2025-08-12 Cunhang Fan , Sheng Zhang , Jingjing Zhang , Enrui Liu , Xinhui Li , Gangming Zhao , Zhao Lv

We investigate the effects of the regularization parameter for the norm () and penalty parameter () in the alternating direction method of multipliers (ADMM) on the quality of restored medical images. Simulation studies are performed using…

Medical Physics · Physics 2021-07-06 Kenya Murase

Distributed phased arrays based multiple-input multiple-output (DPA-MIMO) is a recently proposed highly reconfigurable architecture enabling both spatial multiplexing and beamforming in millimeter-wave (mmWave) systems. In this work, we…

Information Theory · Computer Science 2019-08-27 Yu Zhang , Yiming Huo , Jinlong Zhan , Dongming Wang , Xiaodai Dong , Xiaohu You

In this paper, a generic extension of variational mode decomposition (VMD) algorithm for multivariate or multichannel data sets is presented. We first define a model for multivariate modulated oscillations that is based on the presence of a…

Signal Processing · Electrical Eng. & Systems 2020-01-08 Naveed ur Rehman , Hania Aftab

For frequency division duplex channels, a simple pilot loop-back procedure has been proposed that allows the estimation of the UL & DL channels at an antenna array without relying on any digital signal processing at the terminal side. For…

Information Theory · Computer Science 2017-11-22 Stefan Wesemann , Thomas L. Marzetta

A method for mathematical treatment is considered for experimental data from pulsed time-of-flight spectrometers, whose response to measurement-initiating pulses is represented by their convolution with the pulse response of the system. The…

General Physics · Physics 2018-12-11 Andrey V. Novikov-Borodin

This paper addresses the problem of under-determinded speech source separation from multichannel microphone singals, i.e. the convolutive mixtures of multiple sources. The time-domain signals are first transformed to the short-time Fourier…

Sound · Computer Science 2019-04-11 Xiaofei Li , Laurent Girin , Radu Horaud

Articulatory acoustic inversion aims to reconstruct the complete geometry of the vocal tract from the speech signal. In this paper, we present a comparative study of several levels of phonetic segmentation accuracy, together with a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-13 Sofiane Azzouz , Pierre-André Vuissoz , Yves Laprie

The short-time Fourier transform (STFT) is widely used for analyzing non-stationary signals. However, its performance is highly sensitive to its parameters, and manual or heuristic tuning often yields suboptimal results. To overcome this…

Sound · Computer Science 2025-06-27 Maxime Leiber , Yosra Marnissi , Axel Barrau , Sylvain Meignen , Laurent Massoulié

In many applications in compressed sensing, the measurement matrix is a Fourier matrix, i.e., it measures the Fourier transform of the underlying signal at some specified `base' frequencies $\{u_i\}_{i=1}^M$, where $M$ is the number of…

Information Theory · Computer Science 2018-02-09 Eeshan Malhotra , Himanshu Pandotra , Ajit Rajwade , Karthik S. Gurumoorthy

This paper is concerned with the numerical simulation of three dimensional time-dependent inverse source problems of acoustic waves. The reconstructions of both multiple stationary point sources and a moving point source are considered. The…

Numerical Analysis · Mathematics 2020-08-26 Bo Chen , Yukun Guo , Fuming Ma , Yao Sun

Most state-of-the-art Text-to-Speech systems use the mel-spectrogram as an intermediate representation, to decompose the task into acoustic modelling and waveform generation. A mel-spectrogram is extracted from the waveform by a simple,…

Acoustic-to-articulatory inversion (AAI) methods estimate articulatory movements from the acoustic speech signal, which can be useful in several tasks such as speech recognition, synthesis, talking heads and language tutoring. Most earlier…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-06 Tamás Gábor Csapó