Related papers: A study of vowel nasalization using instantaneous …
Auditory processing difficulties involve challenges in understanding speech in noisy environments despite normal hearing. However, the neural mechanisms remain unclear, and standardized diagnostic criteria are lacking. This study examined…
Acoustic scene classification (ASC) aims to identify the type of scene (environment) in which a given audio signal is recorded. The log-mel feature and convolutional neural network (CNN) have recently become the most popular time-frequency…
Here we present a study of stochastic resonance in an extended FitzHugh-Nagumo system with a field dependent activator diffusion. We show that the system response (here measured through the output signal-to-noise ratio) is enhanced due to…
This paper describes an original experimental procedure to measure the mechanical interaction between the tongue and teeth and palate during speech production. It consists in using edentulous people as subjects and to insert pressure…
Synchronization occurs ubiquitously in nature and science. The synchronization regions generally broaden monotonically with the strength of the forcing, thereby featuring a tongue-like shape in parameter space, known as Arnold's tongue.…
A large eddy simulation is performed to study secondary tones generated by a NACA0012 airfoil at angle of attack of $\alpha = 3^{\circ}$ with freestream Mach number of $M_{\infty} = 0.3$ and Reynolds number of $Re = 5 \times 10^4$. Laminar…
Whistled speech is a little studied local use of language shaped by several cultures of the world either for distant dialogues or for rendering traditional songs. This practice consists of an emulation of the voice thanks to a simple…
Cooperative effects of periodic force and noise in globally Cooperative effects of periodic force and noise in globally coupled systems are studied using a nonlinear diffusion equation for the number density. The amplitude of the order…
Synchronization and resonance on networks are some of the most remarkable collective dynamical phenomena. The network topology, or the nature and distribution of the connections within an ensemble of coupled oscillators, plays a crucial…
Speaker diarization is a task to label audio or video recordings with classes that correspond to speaker identity, or in short, a task to identify "who spoke when". In the early years, speaker diarization algorithms were developed for…
Respiratory diseases remain major global health challenges, and traditional auscultation is often limited by subjectivity, environmental noise, and inter-clinician variability. This study presents an explainable multimodal deep learning…
Despite achieving satisfactory performance in speaker verification using deep neural networks, variable-duration utterances remain a challenge that threatens the robustness of systems. To deal with this issue, we propose a speaker…
The synchronization of coupled organ pipes represents a nonlinear phenomenon with significant implications for both musical acoustics and nonlinear dynamics. This study investigates the coupling mechanisms governing synchronization,…
The dynamics of an organ pipe's mouth region has been studied by numerical simulations. The investigations presented here were carried out by solving the compressible Navier-Stokes equations under suitable initial and boundary conditions…
Neural audio codecs (NACs), which use neural networks to generate compact audio representations, have garnered interest for their applicability to many downstream tasks -- especially quantized codecs due to their compatibility with large…
Excitability, encountered in numerous fields from biology to neurosciences and optics, is a general phenomenon characterized by an all-or-none response of a system to an external perturbation. When subject to delayed feedback, excitable…
Active speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers. The successful ASD depends on accurate interpretation of short-term and long-term audio and visual information, as well as…
Disentangling speaker and content attributes of a speech signal into separate latent representations followed by decoding the content with an exchanged speaker representation is a popular approach for voice conversion, which can be trained…
Audio-visual automatic speech recognition is a promising approach to robust ASR under noisy conditions. However, up until recently it had been traditionally studied in isolation assuming the video of a single speaking face matches the…
A (diatomic) shape resonance is a metastable state of a pair of colliding atoms quasi-bound by the centrifugal barrier imposed by the angular momentum involved in the collision. The temporary trapping of the atoms' scattering wavefunction…