Related papers: How far are vowel formants from computed vocal tra…
Most studies on speaker verification systems focus on long-duration utterances, which are composed of sufficient phonetic information. However, the performances of these systems are known to degrade when short-duration utterances are…
One of the most difficult speech recognition tasks is accurate recognition of human to human communication. Advances in deep learning over the last few years have produced major speech recognition improvements on the representative…
In this paper, we present an improved model for voicing silent speech, where audio is synthesized from facial electromyography (EMG) signals. To give our model greater flexibility to learn its own input features, we directly use EMG signals…
We survey functional analytic methods for studying subwavelength resonator systems. In particular, rigorous discrete approximations of Helmholtz scattering problems are derived in an asymptotic subwavelength regime. This is achieved by…
Phone level localization of mis-articulation is a key requirement for an automatic articulation error assessment system. A robust phone segmentation technique is essential to aid in real-time assessment of phone level mis-articulations of…
We propose a physical model to predict indirect noise generated by the acceleration of compositional inhomogeneities in nozzles with viscous dissipation (non-isentropic nozzles). First, we derive the quasi-one-dimensional equations from the…
We present a method to separate speech signals from noisy environments in the embedding space of a neural audio codec. We introduce a new training procedure that allows our model to produce structured encodings of audio waveforms given by…
A rigorous mathematical theory is developed to explain the super-resolution phenomenon observed in the experiment by F.Lemoult, M.Fink and G.Lerosey (Acoustic resonators for far-field control of sound on a subwavelength scale, Phys. Rev.…
Speaker verification (SV) utilizing features obtained from models pre-trained via self-supervised learning has recently demonstrated impressive performances. However, these pre-trained models (PTMs) usually have a temporal resolution of 20…
We propose a hybrid Finite Volume (FV) - Spectral Element Method (SEM) for modelling aeroacoustic phenomena based on the Lighthill's acoustic analogy. First the fluid solution is computed employing a FV method. Then, the sound source term…
This paper explores whether considering alternative domain-specific embeddings to calculate the Fr\'echet Audio Distance (FAD) metric can help the FAD to correlate better with perceptual ratings of environmental sounds. We used embeddings…
Many machine learning algorithms represent input data with vector embeddings or discrete codes. When inputs exhibit compositional structure (e.g. objects built from parts or procedures from subroutines), it is natural to ask whether this…
Recent progress in audio source separation lead by deep learning has enabled many neural network models to provide robust solutions to this fundamental estimation problem. In this study, we provide a family of efficient neural network…
Speaker embeddings represent a means to extract representative vectorial representations from a speech signal such that the representation pertains to the speaker identity alone. The embeddings are commonly used to classify and discriminate…
Word embedding or Word2Vec has been successful in offering semantics for text words learned from the context of words. Audio Word2Vec was shown to offer phonetic structures for spoken words (signal segments for words) learned from signals…
Sound-soft fractal screens can scatter acoustic waves even when they have zero surface measure. To solve such scattering problems we make what appears to be the first application of the boundary element method (BEM) where each BEM basis…
This work presents a combined numerical and experimental approach to characterize the macroscopic transport and acoustic behavior of foam materials with a membrane cellular structure. A direct link between the sound absorption behavior of a…
We have studied the first three symmetric out-of-plane flexural resonance modes of a goalpost silicon micro-mechanical device. Measurements have been performed at 4.2K in vacuum, demonstrating high Qs and good linear properties. Numerical…
Voiced segments of speech are assumed to be composed of non-stationary acoustic objects which can be described as stationary response of a non-stationary fundamental drive (FD) process and which are furthermore suited to reconstruct the…
Impressive progress in neural network-based single-channel speech source separation has been made in recent years. But those improvements have been mostly reported on anechoic data, a situation that is hardly met in practice. Taking the…