Related papers: Comparison of aerosol emissions during specific sp…
Phonetics is the scientific field concerned with the study of how speech is produced, heard and perceived. It abounds with data, such as acoustic speech recordings, neuroimaging data, or articulatory data. In this paper, we provide an…
Context. Understanding the role of fragmentation is one of the most important current questions of star formation. To better understand the process of star and cluster formation, we need to study in detail the physical structure and…
We investigate the effect of speaker localization on the performance of speech recognition systems in a multispeaker, multichannel environment. Given the speaker location information, speech separation is performed in three stages. In the…
We propose a low-dimensional modeling approach to simulate the dynamics, acoustic emissions and interactions of cavitation bubbles, based on a quasi-acoustic assumption. This quasi-acoustic assumption accounts for the compressibility of the…
The formation of planetesimals requires the growth of dust particles through collisions. Micron-sized particles must grow by many orders of magnitude in mass. In order to understand and model the processes during this growth, the mechanical…
Human spoken language has long been the subject of scientific investigation, particularly with regard to the mechanisms underpinning speech production. Likewise, the study of animal communications has a substantial literature, with many…
Acoustic droplet vaporization denotes the phase-change of micron- and sub-micron-sized droplets upon the application of high-amplitude ultrasound. The asymmetric collapse of the incepted vapor bubbles within the droplets can give rise to…
As the burden of respiratory diseases continues to fall on society worldwide, this paper proposes a high-quality and reliable dataset of human sounds for studying respiratory illnesses, including pneumonia and COVID-19. It consists of…
Speaker identity plays a significant role in human communication and is being increasingly used in societal applications, many through advances in machine learning. Speaker identity perception is an essential cognitive phenomenon that can…
The transport of aerosol discharge in the form of a passive scalar or tracer discharged from a single cough of a patient in a ventilated mock hospital isolation room is investigated via computational fluid dynamics (CFD). Healthcare worker…
It was recently demonstrated that laser filamentation was able to generate an optically transparent channel through cloud and fog for free-space optical communications applications. However, no quantitative measurement of the interaction…
The evolving speech processing landscape is increasingly focused on complex scenarios like meetings or cocktail parties with multiple simultaneous speakers and far-field conditions. Existing methodologies for addressing these challenges…
This paper aims to study the effect of room acoustics and phonemes on the perception of loudness of one's own voice (autophonic loudness) for a group of trained singers. For a set of five phonemes, 20 singers vocalized over several…
Audio-visual automatic speech recognition is a promising approach to robust ASR under noisy conditions. However, up until recently it had been traditionally studied in isolation assuming the video of a single speaking face matches the…
In mobile speech communication applications, wind noise can lead to a severe reduction of speech quality and intelligibility. Since the performance of speech enhancement algorithms using acoustic microphones tends to substantially degrade…
This paper presents a novel approach for detecting mispronunciations by analyzing deviations between a user's original speech and their voice-cloned counterpart with corrected pronunciation. We hypothesize that regions with maximal acoustic…
Speech recognition (ASR) and speaker diarization (SD) models have traditionally been trained separately to produce rich conversation transcripts with speaker labels. Recent advances have shown that joint ASR and SD models can learn to…
This paper introduces the Voices Obscured In Complex Environmental Settings (VOICES) corpus, a freely available dataset under Creative Commons BY 4.0. This dataset will promote speech and signal processing research of speech recorded by…
This work unveils the enigmatic link between phonemes and facial features. Traditional studies on voice-face correlations typically involve using a long period of voice input, including generating face images from voices and reconstructing…
Speech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generated by several sources. Traditionally, these tasks have been…