Related papers: A comparative study of two-dimensional vocal tract…
The goal of this paper is to generate realistic audio with a lightweight and fast diffusion-based vocoder, named FreGrad. Our framework consists of the following three key components: (1) We employ discrete wavelet transform that decomposes…
Vocal fold (VF) motion is fundamental to voice production and diagnosis in speech and health sciences. The motion is a consequence of air flow interacting with elastic vocal fold structures. Motivated by existing lumped mass models and…
Although high-fidelity speech can be obtained for intralingual speech synthesis, cross-lingual text-to-speech (CTTS) is still far from satisfactory as it is difficult to accurately retain the speaker timbres(i.e. speaker similarity) and…
Efficient and accurate numerical simulation of 3D acoustic wave propagation in heterogeneous media plays an important role in the success of seismic full waveform inversion (FWI) problem. In this work, we employed the combined scheme and…
In the process of recording, storage and transmission of time-domain audio signals, errors may be introduced that are difficult to correct in an unsupervised way. Here, we train a convolutional deep neural network to re-synthesize input…
Audio-Visual Segmentation (AVS) aims to segment sound-producing objects in video frames based on the associated audio signal. Prevailing AVS methods typically adopt an audio-centric Transformer architecture, where object queries are derived…
Speech separation has been very successful with deep learning techniques. Substantial effort has been reported based on approaches over spectrogram, which is well known as the standard time-and-frequency cross-domain representation for…
This article presents the use of advanced tools applied to the design of devices that can solve specific acoustic problems, improving the already existing devices based on classic technologies. Specifically, we have used two different…
In many applications, and in particular in seismology, realistic propagation media disperse and attenuate waves. This dissipative behavior can be taken into account by using a viscoacoustic propagation model, which incorporates a complex…
Diadochokinetic speech tasks (DDK), in which participants repeatedly produce syllables, are commonly used as part of the assessment of speech motor impairments. These studies rely on manual analyses that are time-intensive, subjective, and…
Diffusing Alpha-emitters Radiation Therapy (DaRT) is a new method for treating solid tumors using alpha particles. Unlike conventional radiotherapy, where the physical models for dose calculations are known and routinely used, for DaRT a…
Video deblurring is still an unsolved problem due to the challenging spatio-temporal modeling process. While existing convolutional neural network-based methods show a limited capacity for effective spatial and temporal modeling for video…
In text-to-speech (TTS) and voice conversion (VC), acoustic features, such as mel spectrograms, are typically used as synthesis or conversion targets owing to their compactness and ease of learning. However, because the ultimate goal is to…
Text-to-speech (TTS) acoustic models map linguistic features into an acoustic representation out of which an audible waveform is generated. The latest and most natural TTS systems build a direct mapping between linguistic and waveform…
At millimeter-wave frequencies, diffuse scattering from rough surfaces is an important propagation mechanism. Including this mechanism in radio propagation modeling tools,such as ray-tracing, is a key step towards realizing accurate…
Training diffusion models for audiovisual sequences allows for a range of generation tasks by learning conditional distributions of various input-output combinations of the two modalities. Nevertheless, this strategy often requires training…
Traditional sound diffusers are quasi-random phase gratings attached to reflecting surfaces whose purpose is to augment the spatiotemporal incoherence of the acoustic field scattered from reflective surfaces. This configuration allows one…
Speech to speech translation (S2ST) is a transformative technology that bridges global communication gaps, enabling real time multilingual interactions in diplomacy, tourism, and international trade. Our review examines the evolution of…
This study proposes a framework for incorporating wavenumber-domain acoustic reflection coefficients into sound field analysis to characterize direction-dependent material reflection and scattering phenomena. The reflection coefficient is…
Temporal volume images with 3D+t (4D) information are often used in medical imaging to statistically analyze temporal dynamics or capture disease progression. Although deep-learning-based generative models for natural images have been…