Related papers: Modal locking between vocal fold and vocal tract o…
This work addresses friction-induced modal interactions in jointed structures, and their effects on the passive mitigation of vibrations by means of friction damping. Under the condition of (nearly) commensurable natural frequencies, the…
We investigate bubble deformations in an homogeneous and isotropic turbulent flow by means of direct numerical simulations of a single bubble in turbulence. We examine interface deformations by decomposing the local radius into the…
The complex behavior of many natural and engineered systems emerges from the interaction of a small number of effective degrees of freedom. Discovering the physical basis of the interactions between these degrees of freedom directly from…
Text does not fully specify the spoken form, so text-to-speech models must be able to learn from speech data that vary in ways not explained by the corresponding text. One way to reduce the amount of unexplained variation in training data…
Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues to track the target…
Vocal feedback (e.g., `mhm', `yeah', `okay') is an important component of spoken dialogue and is crucial to ensuring common ground in conversational systems. The exact meaning of such feedback is conveyed through both lexical and prosodic…
The ubiquitous phenomenon of synchronization is inherently characteristic of dynamical dissipative non-linear systems. In particular, synchronization has been theoretically and experimentally demonstrated for exciton-polariton condensates…
This paper proposes ESTVocoder, a novel excitation-spectral-transformed neural vocoder within the framework of source-filter theory. The ESTVocoder transforms the amplitude and phase spectra of the excitation into the corresponding speech…
Cross-lingual alignment in pretrained language models enables knowledge transfer across languages. Similar alignment has been reported in Whisper-style speech encoders, based on spoken translation retrieval using representational…
The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how we work with audio. In this work, we make an initial attempt…
Experimental results and their interpretations are presented on the nonlinear acoustic effects of multiple scattered elastic waves in unconsolidated granular media. Short wave packets with a central frequency higher than the so-called…
We present a physics-informed voiced backend renderer for singing-voice synthesis. Given synthetic single-channel audio and a fund-amental--frequency trajectory, we train a time-domain Webster model as a physics-informed neural network to…
Speech foundation models have demonstrated exceptional capabilities in speech-related tasks. Nevertheless, these models often struggle with non-verbal audio data, such as vocalizations, baby crying, etc., which are critical for various…
We introduce a neural auto-encoder that transforms the musical dynamic in recordings of singing voice via changes in voice level. Since most recordings of singing voice are not annotated with voice level we propose a means to estimate the…
A sound synthesis model for woodwind instruments is developed using modal decomposition of the input impedance, accounting for viscothermal losses as well as localized nonlinear losses at the end of the resonator. To extend the definition…
Despite important progress, conversational systems often generate dialogues that sound unnatural to humans. We conjecture that the reason lies in their different training and testing conditions: agents are trained in a controlled "lab"…
The effect of electron-phonon coupling on the current noise in a molecular junction is investigated within a simple model. The model comprises a 1-level bridge representing a molecular level that connects between two free electron…
Current speech production systems predominantly rely on large transformer models that operate as black boxes, providing little interpretability or grounding in the physical mechanisms of human speech. We address this limitation by proposing…
Assessment of voice signals has long been performed with the assumption of periodicity as this facilitates analysis. Near periodicity of normal voice signals makes short-time harmonic modeling an appealing choice to extract vocal feature…
We use linear stability analysis and hybrid lattice Boltzmann simulations to study the dynamical behaviour of an active nematic confined in a channel made of viscoelastic material. We find that the quiescent, ordered active nematic is…