English
Related papers

Related papers: On the Use of a Spectral Glottal Model for the Sou…

200 papers

Mainstream deep learning-based dysarthric speech detection approaches typically rely on processing the magnitude spectrum of the short-time Fourier transform of input signals, while ignoring the phase spectrum. Although considerable insight…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-25 Parvaneh Janbakhshi , Ina Kodrasi

This paper proposes a new loss using short-time Fourier transform (STFT) spectra for the aim of training a high-performance neural speech waveform model that predicts raw continuous speech waveform samples directly. Not only amplitude…

Audio and Speech Processing · Electrical Eng. & Systems 2018-10-31 Shinji Takaki , Toru Nakashika , Xin Wang , Junichi Yamagishi

Short-time Fourier transform (STFT) is used as the front end of many popular successful monaural speech separation methods, such as deep clustering (DPCL), permutation invariant training (PIT) and their various variants. Since the frequency…

Sound · Computer Science 2019-02-05 Ziqiang Shi , Huibin Lin , Liu Liu , Rujie Liu , Jiqing Han

Phase retrieval is a problem encountered not only in speech and audio processing, but in many other fields such as optics. Iterative algorithms based on non-convex set projections are effective and frequently used for retrieving the phase…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-10 Tal Peer , Simon Welker , Timo Gerkmann

Articulatory features can provide interpretable and flexible controls for the synthesis of human vocalizations by allowing the user to directly modify parameters like vocal strain or lip position. To make this manipulation through…

Sound · Computer Science 2023-07-11 David Südholt , Mateo Cámara , Zhiyuan Xu , Joshua D. Reiss

We present a modification to the spectrum differential based direct waveform modification for voice conversion (DIFFVC) so that it can be directly applied as a waveform generation module to voice conversion models. The recently proposed…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-30 Wen-Chin Huang , Yi-Chiao Wu , Kazuhiro Kobayashi , Yu-Huai Peng , Hsin-Te Hwang , Patrick Lumban Tobing , Yu Tsao , Hsin-Min Wang , Tomoki Toda

We present the generalized iterative residual fitting (IRF) for the computation of the spherical harmonic transform (SHT) of band-limited signals on the sphere. The proposed method is based on the partitioning of the subspace of…

Information Theory · Computer Science 2017-09-11 Usama Elahi , Zubair Khalid , Rodney A. Kennedy , Jason D. McEwen

This paper contributes to the understanding of vocal folds oscillation during phonation. In order to test theoretical models of phonation, a new experimental set-up using a deformable vocal folds replica is presented. The replica is shown…

Classical Physics · Physics 2007-10-24 Nicolas Ruty , Annemie Van Hirtum , Xavier Pelorson , Ines Lopez-Arteaga , Avraham Hirschberg

The modeling of speech production often relies on a source-filter approach. Although methods parameterizing the filter have nowadays reached a certain maturity, there is still a lot to be gained for several speech processing applications in…

Sound · Computer Science 2020-01-07 Thomas Drugman , Thierry Dutoit

Reassigned spectrograms have shown advantages in precise formant measuring and inter-speaker differentiation. However, reassigned spectrograms suffer from their inability to visualize larger amounts of data in an easily comprehensible and…

Signal Processing · Electrical Eng. & Systems 2025-09-25 Gabriel J. Griswold , Mark A. Griswold

From a machine learning perspective, the human ability localize sounds can be modeled as a non-parametric and non-linear regression problem between binaural spectral features of sound received at the ears (input) and their sound-source…

Sound · Computer Science 2015-02-12 Yuancheng Luo , Dmitry N. Zotkin , Ramani Duraiswami

We present TVF (Time-Varying Filtering), a low-latency speech enhancement model with 1 million parameters. Combining the interpretability of Digital Signal Processing (DSP) with the adaptability of deep learning, TVF bridges the gap between…

Sound · Computer Science 2026-03-04 Riccardo Rota , Kiril Ratmanski , Jozef Coldenhoff , Milos Cernak

The goal of this work is to recover articulatory information from the speech signal by acoustic-to-articulatory inversion. One of the main difficulties with inversion is that the problem is underdetermined and inversion methods generally…

Computation and Language · Computer Science 2007-05-23 Blaise Potard , Yves Laprie

Traditional speech enhancement techniques modify the magnitude of a speech in time-frequency domain, and use the phase of a noisy speech to resynthesize a time domain speech. This work proposes a complex-valued Gaussian process latent…

Sound · Computer Science 2017-01-02 Sih-Huei Chen , Yuan-Shan Lee , Jia-Ching Wang

The long-tailed distribution is a common phenomenon in the real world. Extracted large scale image datasets inevitably demonstrate the long-tailed property and models trained with imbalanced data can obtain high performance for the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-18 Konstantinos Panagiotis Alexandridis , Shan Luo , Anh Nguyen , Jiankang Deng , Stefanos Zafeiriou

We describe a one-dimensional (1D) unsteady and viscous flow model that is derived from the momentum and mass conservation equations, and to enhance this physics-based model, we use a machine learning approach to determine the unknown…

Fluid Dynamics · Physics 2021-04-07 Zheng Li , Ye Chen , Siyuan Chang , Bernard Rousseau , Haoxiang Luo

This paper introduces WaveGrad 2, a non-autoregressive generative model for text-to-speech synthesis. WaveGrad 2 is trained to estimate the gradient of the log conditional density of the waveform given a phoneme sequence. The model takes an…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-22 Nanxin Chen , Yu Zhang , Heiga Zen , Ron J. Weiss , Mohammad Norouzi , Najim Dehak , William Chan

In this paper, we investigate the application of graph signal processing (GSP) theory in speech enhancement. We first propose a set of shift operators to construct graph speech signals, and then analyze their spectrum in the graph Fourier…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-15 Xue Yan , Zhen Yang , Tingting Wang , Haiyan Guo

This paper addresses the problem of multichannel online dereverberation. The proposed method is carried out in the short-time Fourier transform (STFT) domain, and for each frequency band independently. In the STFT domain, the time-domain…

Sound · Computer Science 2020-11-10 Xiaofei Li , Laurent Girin , Sharon Gannot , Radu Horaud

Diffusion models have recently been shown to be relevant for high-quality speech generation. Most work has been focused on generating spectrograms, and as such, they further require a subsequent model to convert the spectrogram to a…

Sound · Computer Science 2024-03-12 Roi Benita , Michael Elad , Joseph Keshet