English
Related papers

Related papers: On the Use of a Spectral Glottal Model for the Sou…

200 papers

For most of the state-of-the-art speech enhancement techniques, a spectrogram is usually preferred than the respective time-domain raw data since it reveals more compact presentation together with conspicuous temporal information over a…

Sound · Computer Science 2016-08-24 Syu-Siang Wang , Alan Chern , Yu Tsao , Jeih-weih Hung , Xugang Lu , Ying-Hui Lai , Borching Su

Video-assisted transoral tracheal intubation (TI) necessitates using an endoscope that helps the physician insert a tracheal tube into the glottis instead of the esophagus. The growing trend of robotic-assisted TI would require a medical…

Artificial Intelligence · Computer Science 2023-07-31 Guankun Wang , Tian-Ao Ren , Jiewen Lai , Long Bai , Hongliang Ren

Flow Matching (FM) is a simulation-free method for learning a continuous and invertible flow to interpolate between two distributions, and in particular to generate data from noise. Inspired by the variational nature of the diffusion…

Machine Learning · Statistics 2025-07-14 Chen Xu , Xiuyuan Cheng , Yao Xie

We describe a new algorithm to solve a particular phase retrieval problem, that has wide applications in audio processing: the reconstruction of a function from its scalogram, that is from the modulus of its wavelet transform. It is a…

Optimization and Control · Mathematics 2017-04-11 Irène Waldspurger

Flow maps enable high-quality image generation in a single forward pass. However, unlike iterative diffusion models, their lack of an explicit sampling trajectory impedes incorporating external constraints for conditional generation and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Abbas Mammadov , So Takao , Bohan Chen , Ricardo Baptista , Morteza Mardani , Yee Whye Teh , Julius Berner

Acoustic-to-articulatory inversion (AAI) aims to estimate the parameters of articulators from speech audio. There are two common challenges in AAI, which are the limited data and the unsatisfactory performance in speaker independent…

Sound · Computer Science 2023-02-28 Jianrong Wang , Jinyu Liu , Li Liu , Xuewei Li , Mei Yu , Jie Gao , Qiang Fang

The performance of speaker-related systems usually degrades heavily in practical applications largely due to the presence of background noise. To improve the robustness of such systems in unknown noisy environments, this paper proposes a…

Sound · Computer Science 2018-05-04 Siyang Song , Shuimei Zhang , Björn Schuller , Linlin Shen , Michel Valstar

Audio source separation is often achieved by estimating the magnitude spectrogram of each source, and then applying a phase recovery (or spectrogram inversion) algorithm to retrieve time-domain signals. Typically, spectrogram inversion is…

Sound · Computer Science 2023-07-03 Paul Magron , Tuomas Virtanen

A non-iterative method for the construction of the Short-Time Fourier Transform (STFT) phase from the magnitude is presented. The method is based on the direct relationship between the partial derivatives of the phase and the logarithm of…

Sound · Computer Science 2019-03-27 Zdeněk Průša , Peter Balazs , Peter L. Søndergaard

This study investigates phase reconstruction for deep learning based monaural talker-independent speaker separation in the short-time Fourier transform (STFT) domain. The key observation is that, for a mixture of two sources, with their…

Sound · Computer Science 2018-11-26 Zhong-Qiu Wang , Ke Tan , DeLiang Wang

Various information factors are blended in speech signals, which forms the primary difficulty for most speech information processing tasks. An intuitive idea is to factorize speech signal into individual information factors (e.g., phonetic…

Sound · Computer Science 2020-10-28 Haoran Sun , Lantian Li , Yunqi Cai , Yang Zhang , Thomas Fang Zheng , Dong Wang

Gaussian process (GP) audio source separation is a time-domain approach that circumvents the inherent phase approximation issue of spectrogram based methods. Furthermore, through its kernel, GPs elegantly incorporate prior knowledge about…

Audio and Speech Processing · Electrical Eng. & Systems 2018-11-22 Pablo A. Alvarado , Mauricio A. Álvarez , Dan Stowell

Conventional NMF methods for source separation factorize the matrix of spectral magnitudes. Spectral Phase is not included in the decomposition process of these methods. However, phase of the speech mixture is generally used in…

Sound · Computer Science 2014-11-26 Chaitanya Ahuja , Karan Nathwani , Rajesh M. Hegde

Reverberation is damaging to both the quality and the intelligibility of a speech signal. We propose a novel single-channel method of dereverberation based on a linear filter in the Short Time Fourier Transform domain. Each enhanced frame…

Sound · Computer Science 2015-09-25 Richard Stanton , Mike Brookes

Generative adversarial network (GAN) models can synthesize highquality audio signals while ensuring fast sample generation. However, they are difficult to train and are prone to several issues including mode collapse and divergence. In this…

Sound · Computer Science 2024-02-06 Teysir Baoueb , Haocheng Liu , Mathieu Fontaine , Jonathan Le Roux , Gael Richard

Diffusion models have recently shown promising results for difficult enhancement tasks such as the conditional and unconditional restoration of natural images and audio signals. In this work, we explore the possibility of leveraging a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-24 Hao Yen , François G. Germain , Gordon Wichern , Jonathan Le Roux

Acoustic beamforming models typically assume wide-sense stationarity of speech signals within short time frames. However, voiced speech is better modeled as a cyclostationary (CS) process, a random process whose mean and autocorrelation are…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-16 Giovanni Bologni , Richard Heusdens , Richard C. Hendriks

Insect population numbers and biodiversity have been rapidly declining with time, and monitoring these trends has become increasingly important for conservation measures to be effectively implemented. But monitoring methods are often…

Sound · Computer Science 2024-02-01 Marius Faiß , Dan Stowell

Speech produced by human vocal apparatus conveys substantial non-semantic information including the gender of the speaker, voice quality, affective state, abnormalities in the vocal apparatus etc. Such information is attributed to the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-16 Prathosh A. P. , Varun Srivastava , Mayank Mishra

We address speech enhancement based on variational autoencoders, which involves learning a speech prior distribution in the time-frequency (TF) domain. A zero-mean complex-valued Gaussian distribution is usually assumed for the generative…

Sound · Computer Science 2023-10-27 Ali Golmakani , Mostafa Sadeghi , Xavier Alameda-Pineda , Romain Serizel