Related papers: Modal locking between vocal fold and vocal tract o…
Vocal fold (VF) motion is a fundamental process in voice production, and is also a challenging problem for direct numerical computation because the VF dynamics depend on nonlinear coupling of air flow with the response of elastic channels…
We describe an arrangement for simultaneous recording of speech and geometry of vocal tract in patients undergoing surgery involving this area. Experimental design is considered from an articulatory phonetic point of view. The speech and…
Using an artificial mouth with an accurate pressure control, the onset of the pressure oscillations inside the mouthpiece of a simplified clarinet is studied experimentally. Two time profiles are used for the blowing pressure: in a first…
Sound morphing is the process of gradually and smoothly transforming one sound into another to generate novel and perceptually hybrid sounds that simultaneously resemble both. Recently, diffusion-based text-to-audio models have produced…
All previous methods for audio-driven talking head generation assume the input audio to be clean with a neutral tone. As we show empirically, one can easily break these systems by simply adding certain background noise to the utterance or…
Several experiments have been performed to investigate the mechanical vibrations associated with trachea and larynx when Italian vowels are emitted. The mechanical measurements have been made by using two laser Doppler vibrometers (based on…
What do deep neural speech models know about phonology? Existing work has examined the encoding of individual linguistic units such as phonemes in these models. Here we investigate interactions between units. Inspired by classic experiments…
Spectro-temporal dynamics of consonant-vowel (CV) transition regions are considered to provide robust cues related to articulation. In this work, we propose an objective measure of precise articulation, dubbed the objective articulation…
We present a method for analyzing the phase noise of oscillators based on feedback driven high quality factor resonators. Our approach is to derive the phase drift of the oscillator by projecting the stochastic oscillator dynamics onto a…
The recent wave of audio foundation models (FMs) could provide new capabilities for conversational modeling. However, there have been limited efforts to evaluate these audio FMs comprehensively on their ability to have natural and…
Social interaction dynamics are a special type of group interactions that play a large part in our everyday lives. They dictate how and with whom a certain individual will interact. One of such interactions can be termed "avoidance…
Multiple studies in the past have shown that there is a strong correlation between human vocal characteristics and facial features. However, existing approaches generate faces simply from voice, without exploring the set of features that…
Rapid advances in speech synthesis and audio editing have made realistic forgeries increasingly accessible, yet existing detection methods remain vulnerable to tampering or depend on visual/wearable sensors. In this paper, we present…
We propose a low-dimensional modeling approach to simulate the dynamics, acoustic emissions and interactions of cavitation bubbles, based on a quasi-acoustic assumption. This quasi-acoustic assumption accounts for the compressibility of the…
A model for the acoustic production of gravitational waves at a first order phase transition is presented. The source of gravitational radiation is the sound waves generated by the explosive growth of bubbles of the stable phase. The model…
This paper presents a study of the forced acoustical response of an open cavity from the perspective of modal expansion. Based on the coupled mode theory, it is shown that the sound pressure distribution of an open cavity excited by a point…
Speech is produced through the coordination of vocal tract constricting organs: lips, tongue, velum, and glottis. Previous works developed Speech Inversion (SI) systems to recover acoustic-to-articulatory mappings for lip and tongue…
Phonetic production bias is the external force most commonly invoked in computational models of sound change, despite the fact that it is not responsible for all, or even most, sound changes. Furthermore, the existence of production bias…
We explore the acoustic phonon-based interaction between two neighboring coplanar circuits containing semiconductor quantum point contacts in a perpendicular magnetic field B. In a drag-type experiment, a current flowing in one of the…
We propose the use of a self-oscillating dynamical system --the pre-Galileian clock equation-- for modeling the laryngeal tone. The parameters are shown to be the minimal control needed for generating the prosody of the human speech. Based…