Related papers: A study of vowel nasalization using instantaneous …
Articulatory-to-acoustic (forward) mapping is a technique to predict speech using various articulatory acquisition techniques (e.g. ultrasound tongue imaging, lip video). Real-time MRI (rtMRI) of the vocal tract has not been used before for…
Acoustic-to-articulatory inversion (AAI) methods estimate articulatory movements from the acoustic speech signal, which can be useful in several tasks such as speech recognition, synthesis, talking heads and language tutoring. Most earlier…
To investigate how the auditory system processes natural speech, models have been created to relate the electroencephalography (EEG) signal of a person listening to speech to various representations of the speech. Mainly the speech envelope…
Vocal entrainment is a social adaptation mechanism in human interaction, knowledge of which can offer useful insights to an individual's cognitive-behavioral characteristics. We propose a context-aware approach for measuring vocal…
Recent speech enhancement methods based on convolutional neural networks (CNNs) and transformer have been demonstrated to efficaciously capture time-frequency (T-F) information on spectrogram. However, the correlation of each channels of…
The study of sound propagation in a uniform duct having a mean flow has many applications, such as in the design of gas turbines, heating, ventilation and air conditioning ducts, automotive intake and exhaust systems, and in the modeling of…
Recently, several studies have investigated the transcription process associated to specific genetic regulatory networks. In this work, we present a stochastic approach for analyzing the dynamics and effect of negative feedback loops (FBL)…
We analyze the stochastic response of a finite set of globally coupled noisy bistable units driven by rather weak time-periodic forces. We focus on the stochastic resonance and phase frequency synchronization of the collective variable,…
Although diffusion models in text-to-speech have become a popular choice due to their strong generative ability, the intrinsic complexity of sampling from diffusion models harms their efficiency. Alternatively, we propose VoiceFlow, an…
In this paper, a novel architecture for speaker recognition is proposed by cascading speech enhancement and speaker processing. Its aim is to improve speaker recognition performance when speech signals are corrupted by noise. Instead of…
Sudden turn-on of a matter-wave source leads to characteristic oscillations of the probability density which are the hallmark feature of diffraction in time. The apodization of matter waves relies on the use of smooth aperture functions…
In what ways might statistical signals in linguistic input assist with the acquisition of syntax? Here we hypothesize a mechanism called collocational bootstrapping, in which regularities in word co-occurrence patterns can provide cues to…
We describe here an experimental technique based on the acoustic scattering phenomenon allowing the direct probing of the vorticity field in a turbulent flow. Using time-frequency distributions, recently introduced in signal analysis…
Resonance, defined as the oscillation of a system when the temporal frequency of an external stimulus matches a natural frequency of the system, is important in both fundamental physics and applied disciplines. However, the spatial…
There are two types of methods for non-autoregressive text-to-speech models to learn the one-to-many relationship between text and speech effectively. The first one is to use an advanced generative framework such as normalizing flow (NF).…
The historical and geographical spread from older to more modern languages has long been studied by examining textual changes and in terms of changes in phonetic transcriptions. However, it is more difficult to analyze language change from…
The inverse problem of determining the cross-sectional area of a human vocal tract during the utterance of a vowel is considered in terms of the data consisting of the absolute value of sound pressure at the lips. If the upper lip is curved…
Discontinuous shear thickening (DST) in dense suspensions leads to flow instabilities that limit processing in many systems. While high-power ultrasound has been reported to reduce the apparent viscosity of such materials, the origin of…
To investigate how speech is processed in the brain, we can model the relation between features of a natural speech signal and the corresponding recorded electroencephalogram (EEG). Usually, linear models are used in regression tasks.…
Singing techniques are used for expressive vocal performances by employing temporal fluctuations of the timbre, the pitch, and other components of the voice. Their classification is a challenging task, because of mainly two factors: 1) the…