English
Related papers

Related papers: On the Use of a Spectral Glottal Model for the Sou…

200 papers

Source-tract decomposition (or glottal flow estimation) is one of the basic problems of speech processing. For this, several techniques have been proposed in the literature. However studies comparing different approaches are almost…

Sound · Computer Science 2020-01-06 Thomas Drugman , Baris Bozkurt , Thierry Dutoit

This paper introduces GlOttal-flow LPC Filter (GOLF), a novel method for singing voice synthesis (SVS) that exploits the physical characteristics of the human voice using differentiable digital signal processing. GOLF employs a glottal…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-21 Chin-Yun Yu , György Fazekas

Some glottal analysis approaches based upon linear prediction or complex cepstrum approaches have been proved to be effective to estimate glottal source from real speech utterances. We propose a new approach employing both an all-pole…

Sound · Computer Science 2016-12-16 Yiqiao Chen , John N. Gowdy

This paper addresses the problem of estimating the voice source directly from speech waveforms. A novel principle based on Anticausality Dominated Regions (ACDR) is used to estimate the glottal open phase. This technique is compared to two…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-26 Thomas Drugman , Thomas Dubuisson , Alexis Moinet , Nicolas D'Alessandro , Thierry Dutoit

Modeling and estimation of the vocal tract and glottal source parameters of vowels from raw speech can be typically done by using the Auto-Regressive with eXogenous input (ARX) model and Liljencrants-Fant (LF) model with an iteration-based…

Sound · Computer Science 2024-10-08 Kai Lia , Masato Akagia , Yongwei Lib , Masashi Unokia

Automatic detection of voice pathology enables objective assessment and earlier intervention for the diagnosis. This study provides a systematic analysis of glottal source features and investigates their effectiveness in voice pathology…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-18 Sudarsana Reddy Kadiri , Paavo Alku

For audio source separation applications, it is common to estimate the magnitude of the short-time Fourier transform (STFT) of each source. In order to further synthesizing time-domain signals, it is necessary to recover the phase of the…

Sound · Computer Science 2018-02-28 Paul Magron , Roland Badeau , Bertrand David

The great majority of current voice technology applications relies on acoustic features characterizing the vocal tract response, such as the widely used MFCC of LPC parameters. Nonetheless, the airflow passing through the vocal folds, and…

Sound · Computer Science 2020-01-01 Thomas Drugman , Paavo Alku , Abeer Alwan , Bayya Yegnanarayana

Formant synthesis aims to generate speech with controllable formant structures, enabling precise control of vocal resonance and phonetic features. However, while existing formant synthesis approaches enable precise formant manipulation,…

Although diffusion models in text-to-speech have become a popular choice due to their strong generative ability, the intrinsic complexity of sampling from diffusion models harms their efficiency. Alternatively, we propose VoiceFlow, an…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-04 Yiwei Guo , Chenpeng Du , Ziyang Ma , Xie Chen , Kai Yu

The spectrotemporal receptive field (STRF) provides a versatile and integrated, spectral and temporal, functional characterization of single cells in primary auditory cortex (AI). In this paper, we explore the origin of, and relationship…

Neurons and Cognition · Quantitative Biology 2007-05-23 David J. Klein , Jonathan Z. Simon , Didier A. Depireux , Shihab A. Shamma

Homomorphic analysis is a well-known method for the separation of non-linearly combined signals. More particularly, the use of complex cepstrum for source-tract deconvolution has been discussed in various articles. However there exists no…

Sound · Computer Science 2020-01-01 Thomas Drugman , Baris Bozkurt , Thierry Dutoit

This paper addresses the problem of under-determinded speech source separation from multichannel microphone singals, i.e. the convolutive mixtures of multiple sources. The time-domain signals are first transformed to the short-time Fourier…

Sound · Computer Science 2019-04-11 Xiaofei Li , Laurent Girin , Radu Horaud

This paper addresses the problem of speech separation and enhancement from multichannel convolutive and noisy mixtures, \emph{assuming known mixing filters}. We propose to perform the speech separation and enhancement task in the short-time…

Sound · Computer Science 2019-01-31 Xiaofei Li , Laurent Girin , Sharon Gannot , Radu Horaud

The human vocal folds are known to interact with the vocal tract acoustics during voiced speech production; namely a nonlinear source-filter coupling has been observed both by using models and in \emph{in vivo} phonation. These phenomena…

Biological Physics · Physics 2015-11-17 Daniel Aalto , Jarmo Malinen , Martti Vainio

Acoustic-to-articulatory inversion (AAI) methods estimate articulatory movements from the acoustic speech signal, which can be useful in several tasks such as speech recognition, synthesis, talking heads and language tutoring. Most earlier…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-06 Tamás Gábor Csapó

Pitch and Formant frequencies are important features in speech processing applications. The period of the vocal cord's output for vowels is known as the pitch or the fundamental frequency, and formant frequencies are essentially resonance…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-09 Seyedamiryousef Hosseini Goki , Mahdieh Ghazvini , Sajad Hamzenejadi

Head-related transfer functions (HRTFs) are important for immersive audio, and their spatial interpolation has been studied to upsample finite measurements. Recently, neural fields (NFs) which map from sound source direction to HRTF have…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-29 Yoshiki Masuyama , Gordon Wichern , François G. Germain , Zexu Pan , Sameer Khurana , Chiori Hori , Jonathan Le Roux

Articulatory acoustic inversion aims to reconstruct the complete geometry of the vocal tract from the speech signal. In this paper, we present a comparative study of several levels of phonetic segmentation accuracy, together with a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-13 Sofiane Azzouz , Pierre-André Vuissoz , Yves Laprie

This paper introduces a phase-aware probabilistic model for audio source separation. Classical source models in the short-term Fourier transform domain use circularly-symmetric Gaussian or Poisson random variables. This is equivalent to…

Sound · Computer Science 2018-10-02 Paul Magron , Tuomas Virtanen
‹ Prev 1 2 3 10 Next ›