Related papers: How far are vowel formants from computed vocal tra…
Head-related transfer functions (HRTFs) describe the directional filtering of the incoming sound caused by the morphology of a listener's head and pinnae. When an accurate model of a listener's morphology exists, HRTFs can be calculated…
We introduce a novel virtual element method (VEM) for the two dimensional Helmholtz problem endowed with impedance boundary conditions. Local approximation spaces consist of Trefftz functions, i.e., functions belonging to the kernel of the…
Previous works on voice-face matching and voice-guided face synthesis demonstrate strong correlations between voice and face, but mainly rely on coarse semantic cues such as gender, age, and emotion. In this paper, we aim to investigate the…
The scattering and transmission of harmonic acoustic waves at a penetrable material are commonly modelled by a set of Helmholtz equations. This system of partial differential equations can be rewritten into boundary integral equations…
We propose a new method for speaker diarization that can handle overlapping speech with 2+ people. Our method is based on compositional embeddings [1]: Like standard speaker embedding methods such as x-vector [2], compositional embedding…
Distant speech processing is a challenging task, especially when dealing with the cocktail party effect. Sound source separation is thus often required as a preprocessing step prior to speech recognition to improve the signal to distortion…
The Minneart resonance is a low frequency resonance in which the wavelength is much larger than the size of the resonators. It is interesting to study the interaction between two adjacent bubbles when they are brought close together.…
Classroom environments are particularly challenging for children with hearing impairments, where background noise, multiple talkers, and reverberation degrade speech perception. These difficulties are greater for children than adults, yet…
Target speech extraction remains difficult for compact devices because monaural neural models lack spatial evidence and classical beamformers lose resolving power when the microphone aperture is only a few centimetres. We present IsoNet, a…
This paper presents a shape optimisation system to design the shape of an acoustically-hard object in the three-dimensional open space. Boundary element method (BEM) is suitable to analyse such an exterior field. However, the conventional…
Resolvent analysis has demonstrated encouraging results for modeling coherent structures in jets when compared against their data-educed counterparts from high-fidelity large-eddy simulations (LES). We formulate resolvent analysis as an…
The auditory ossicles that are located in the middle ear are the smallest bones in the human body. Their damage will result in hearing loss. It is therefore important to be able to automatically diagnose ossicles' diseases based on Computed…
In this paper, we develop a numerical method for the computation of (quasi-)resonances in spherical symmetric, heterogeneous Helmholtz problems with piecewise smooth refractive index. Our focus lies in resonances very close to the real…
The Glottal Source is an important component of voice as it can be considered as the excitation signal to the voice apparatus. Nowadays, new techniques of speech processing such as speech recognition and speech synthesis use the glottal…
The source separation-based speech enhancement problem with multiple beamforming in reverberant indoor environments is addressed in this paper. We propose that more generic solutions should cope with time-varying dynamic scenarios with…
In traditional studies on language evolution, scholars often emphasize the importance of sound laws and sound correspondences for phylogenetic inference of language family trees. However, to date, computational approaches have typically not…
In this paper, we compute the band structure of one- and two-dimensional phononic composites using the extended finite element method (X-FEM) on structured higher-order (spectral) finite element meshes. On using partition-of-unity…
Extracting features from the speech is the most critical process in speech signal processing. Mel Frequency Cepstral Coefficients (MFCC) are the most widely used features in the majority of the speaker and speech recognition applications,…
This paper presents a subjective study conducted on the perception of salient auditory attributes depending on the listener's position and head orientations in an enclosed space. Two elicitation experiments were carried out using the…
It is shown that in a rotating compressible fluid the resonant frequencies (measured in a system of reference rotating together with the medium) for the azimuthally running acoustic waves are split into two components. The received results…