Related papers: How far are vowel formants from computed vocal tra…
This paper presents predictions of the consequences of tongue surgery on speech production. For this purpose, a 3D finite element model of the tongue is used that represents this articulator as a deformable structure in which tongue muscles…
The Generation and propagation of the human voice is studied in two-dimensions using a full-body domain, using direct numerical simulation. The fluid/air in the vocal tract is modeled as a compressible and viscous fluid interacting with the…
The inverse problem of determining the cross-sectional area of a human vocal tract during the utterance of a vowel is considered in terms of the data consisting of the absolute value of sound pressure at the lips. If the upper lip is curved…
The three-dimensional reconstruction of vocal folds in medicine usually involves endoscopy and an approach to extract depth information like structured light or stereo matching of images. The resulting mesh can accurately represent the…
We present a physics-informed voiced backend renderer for singing-voice synthesis. Given synthetic single-channel audio and a fund-amental--frequency trajectory, we train a time-domain Webster model as a physics-informed neural network to…
Whispered speech is characterised by a noise-like excitation that results in the lack of fundamental frequency. Considering that prosodic phenomena such as intonation are perceived through f0 variation, the perception of whispered prosody…
We analyze the Helmholtz equation in a complex domain. A sound absorbing structure at a part of the boundary is modelled by a periodic geometry with periodicity $\varepsilon>0$. A resonator volume of thickness $\varepsilon$ is connected…
Introduction Speech is an integral component of human communication, requiring the coordinated efforts of various organs to produce sound (Titze & Alipour, 2006). The glottis region, a key player in voice production, assumes a crucial role…
While Word2Vec represents words (in text) as vectors carrying semantic information, audio Word2Vec was shown to be able to represent signal segments of spoken words as vectors carrying phonetic structure information. Audio Word2Vec can be…
The study of aerosols and droplets emitted from the oral cavity has become increasingly important throughout the COVID-19 pandemic. Studies show particulates emitted while speaking were generally much smaller compared to coughing or…
Changing the vocal tract shape is one of the techniques which can be used by the players of wind instruments to modify the quality of the sound. It has been intensely studied in the case of reed instruments but has received only little…
Vocal tract configurations play a vital role in generating distinguishable speech sounds, by modulating the airflow and creating different resonant cavities in speech production. They contain abundant information that can be utilized to…
Formants are the spectral maxima that result from acoustic resonances of the human vocal tract, and their accurate estimation is among the most fundamental speech processing problems. Recent work has been shown that those frequencies can…
In this work we have developed a technique for the measurement of the resonance curve of Helmholtz resonators as a function of filling with beads and sands of different sizes, and water as the reference. Our measurements allowed us to…
We can estimate the size of the speakers based on their speech sounds alone. We had proposed an auditory computational theory of the Stabilised Wavelet-Mellin Transform (SWMT), which segregates information about the size and shape of the…
A key barrier to making phonetic studies scalable and replicable is the need to rely on subjective, manual annotation. To help meet this challenge, a machine learning algorithm was developed for automatic measurement of a widely used…
The way infants use auditory cues to learn to speak despite the acoustic mismatch of their vocal apparatus is a hot topic of scientific debate. The simulation of early vocal learning using articulatory speech synthesis offers a way towards…
In this study, the performance of Helmholtz resonators with curved tapered neck extensions was investigated through numerical and experimental methods. A numerical parametric analysis was carried out using the Finite Element Method with…
Understanding speech production both visually and kinematically can inform second language learning system designs, as well as the creation of speaking characters in video games and animations. In this work, we introduce a data-driven…
Spectro-temporal dynamics of consonant-vowel (CV) transition regions are considered to provide robust cues related to articulation. In this work, we propose an objective measure of precise articulation, dubbed the objective articulation…