English
Related papers

Related papers: Vocal Tract Area Estimation by Gradient Descent

200 papers

Disorders of voice production have severe effects on the quality of life of the affected individuals. A simulation approach is used to investigate the cause-effect chain in voice production showing typical characteristics of voice such as…

Sound · Computer Science 2022-07-20 Florian Kraxberger , Andreas Wurzinger , Stefan Schoder

Singing techniques are used for expressive vocal performances by employing temporal fluctuations of the timbre, the pitch, and other components of the voice. Their classification is a challenging task, because of mainly two factors: 1) the…

Sound · Computer Science 2022-06-27 Yuya Yamamoto , Juhan Nam , Hiroko Terasawa

Automatic lyrics to polyphonic audio alignment is a challenging task not only because the vocals are corrupted by background music, but also there is a lack of annotated polyphonic corpus for effective acoustic modeling. In this work, we…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-26 Chitralekha Gupta , Emre Yılmaz , Haizhou Li

Vocal tract configurations play a vital role in generating distinguishable speech sounds, by modulating the airflow and creating different resonant cavities in speech production. They contain abundant information that can be utilized to…

Sound · Computer Science 2018-07-31 Pramit Saha , Praneeth Srungarapu , Sidney Fels

Syllable detection is an important speech analysis task with applications in speech rate estimation, word segmentation, and automatic prosody detection. Based on the well understood acoustic correlates of speech articulation, it has been…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-09 Kamini Sabu , Syomantak Chaudhuri , Preeti Rao , Mahesh Patil

Speech is produced through the coordination of vocal tract constricting organs: lips, tongue, velum, and glottis. Previous works developed Speech Inversion (SI) systems to recover acoustic-to-articulatory mappings for lip and tongue…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-01 Saba Tabatabaee , Suzanne Boyce , Liran Oren , Mark Tiede , Carol Espy-Wilson

In this work, a Bayesian approach to speaker normalization is proposed to compensate for the degradation in performance of a speaker independent speech recognition system. The speaker normalization method proposed herein uses the technique…

Sound · Computer Science 2016-10-20 Dhananjay Ram , Debasis Kundu , Rajesh M. Hegde

Automatic speaker recognition algorithms typically use pre-defined filterbanks, such as Mel-Frequency and Gammatone filterbanks, for characterizing speech audio. However, it has been observed that the features extracted using these…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-14 Anurag Chowdhury , Arun Ross

There are different algorithms for vocal fold pathology diagnosis. These algorithms usually have three stages which are Feature Extraction, Feature Reduction and Classification. While the third stage implies a choice of a variety of machine…

Machine Learning · Computer Science 2013-02-08 Vahid Majidnezhad , Igor Kheidorov

The throat microphone is a body-attached transducer that is worn against the neck. It captures the signals that are transmitted through the vocal folds, along with the buzz tone of the larynx. Due to its skin contact, it is more robust to…

Audio and Speech Processing · Electrical Eng. & Systems 2018-04-18 Mehmet Ali Tugtekin Turan

Some glottal analysis approaches based upon linear prediction or complex cepstrum approaches have been proved to be effective to estimate glottal source from real speech utterances. We propose a new approach employing both an all-pole…

Sound · Computer Science 2016-12-16 Yiqiao Chen , John N. Gowdy

This paper addresses the problem of estimating the voice source directly from speech waveforms. A novel principle based on Anticausality Dominated Regions (ACDR) is used to estimate the glottal open phase. This technique is compared to two…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-26 Thomas Drugman , Thomas Dubuisson , Alexis Moinet , Nicolas D'Alessandro , Thierry Dutoit

Objective: Voice disorders significantly compromise individuals' ability to speak in their daily lives. Without early diagnosis and treatment, these disorders may deteriorate drastically. Thus, automatic classification systems at home are…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-27 Heng-Cheng Kuo , Yu-Peng Hsieh , Huan-Hsin Tseng , Chi-Te Wang , Shih-Hau Fang , Yu Tsao

During voiced speech, the human vocal folds interact with the vocal tract acoustics. The resulting glottal source-resonator coupling has been observed using mathematical and physical models as well as in in vivo phonation. We propose a…

Fluid Dynamics · Physics 2017-03-16 Atte Aalto , Tiina Murtola , Jarmo Malinen , Daniel Aalto , Martti Vainio

Acoustic-to-articulatory inversion (AAI) methods estimate articulatory movements from the acoustic speech signal, which can be useful in several tasks such as speech recognition, synthesis, talking heads and language tutoring. Most earlier…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-06 Tamás Gábor Csapó

In speech processing pipelines, improving the quality and intelligibility of real-world recordings is crucial. While supervised regression is the primary method for speech enhancement, audio tokenization is emerging as a promising…

Sound · Computer Science 2025-07-18 Luca Della Libera , Cem Subakan , Mirco Ravanelli

Audio-to-score alignment is an important pre-processing step for in-depth analysis of classical music. In this paper, we apply novel transposition-invariant audio features to this task. These low-dimensional features represent local pitch…

Sound · Computer Science 2018-07-20 Andreas Arzt , Stefan Lattner

This article develops a general detection theory for speech analysis based on time-varying autoregressive models, which themselves generalize the classical linear predictive speech analysis framework. This theory leads to a computationally…

Applications · Statistics 2011-08-25 Daniel Rudoy , Thomas F. Quatieri , Patrick J. Wolfe

We propose ARTI-6, a compact six-dimensional articulatory speech encoding framework derived from real-time MRI data that captures crucial vocal tract regions including the velum, tongue root, and larynx. ARTI-6 consists of three components:…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-27 Jihwan Lee , Sean Foley , Thanathai Lertpetchpun , Kevin Huang , Yoonjeong Lee , Tiantian Feng , Louis Goldstein , Dani Byrd , Shrikanth Narayanan

This paper describes a human-in-the-loop approach to personalized voice synthesis in the absence of reference speech data from the target speaker. It is intended to help vocally disabled individuals restore their lost voices without…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-27 Yusheng Tian , Junbin Liu , Tan Lee