Related papers: A comparative study of two-dimensional vocal tract…
A conformal dispersive finite-difference time-domain (FDTD) method is developed for the study of one-dimensional (1-D) plasmonic waveguides formed by an array of periodic infinite-long silver cylinders at optical frequencies. The curved…
Multi-resolution spectro-temporal features of a speech signal represent how the brain perceives sounds by tuning cortical cells to different spectral and temporal modulations. These features produce a higher dimensional representation of…
The numerical simulation of acoustic waves in complex 3D media is a key topic in many branches of science, from exploration geophysics to non-destructive testing and medical imaging. With the drastic increase in computing capabilities this…
We propose a multi-channel speech enhancement approach with a novel two-stage feature fusion method and a pre-trained acoustic model in a multi-task learning paradigm. In the first fusion stage, the time-domain and frequency-domain features…
In many industries, including aerospace and defense, waveform analysis is commonly conducted to compute the resonance of physical objects, with the Finite Element Method (FEM) being the standard approach. The Finite Difference Method (FDM)…
Equivocal 3D lesion segmentation exhibits high inter-observer variability. Conventional deterministic models ignore this aleatoric uncertainty, producing over-confident masks that obscure clinical risks. Conversely, while generative methods…
Audio and video are two most common modalities in the mainstream media platforms, e.g., YouTube. To learn from multimodal videos effectively, in this work, we propose a novel audio-video recognition approach termed audio video Transformer,…
For 6-DOF (degrees of freedom) interactive virtual acoustic environments (VAEs), the spatial rendering of diffuse late reverberation in addition to early (specular) reflections is important. In the interest of computational efficiency, the…
Text-to-speech(TTS) has undergone remarkable improvements in performance, particularly with the advent of Denoising Diffusion Probabilistic Models (DDPMs). However, the perceived quality of audio depends not solely on its content, pitch,…
Porous acoustic absorbers have excellent properties in the low-frequency range when positioned in room edges, therefore they are a common method for reducing low-frequency reverberation. However, standard room acoustic simulation methods…
A comprehensive study on the Finite Difference Time Domain (FDTD) numerical modelling of space- and time-varying media is presented. We investigate the dynamic behavior of oblique incidence of both TM and TE electromagnetic fields on…
Voice Type Discrimination (VTD) refers to discrimination between regions in a recording where speech was produced by speakers that are physically within proximity of the recording device ("Live Speech") from speech and other types of audio…
With recent advances of AIGC, video generation have gained a surge of research interest in both academia and industry (e.g., Sora). However, it remains a challenge to produce temporally aligned audio to synchronize the generated video,…
This paper presents the ultra-wideband (UWB) on-body radio channel modelling using a sub-band Finite-Difference Time-Domain (FDTD) method and a model combining the uniform geometrical theory of diffraction (UTD) and ray tracing (RT). In the…
Finite-difference time-domain (FDTD) is an effective algorithm for resolving Maxwell equations directly in time domain. Although FDTD has obtained sufficient development, there still exists some improvement space for it, such as…
This paper presents an audio-visual approach for voice separation which produces state-of-the-art results at a low latency in two scenarios: speech and singing voice. The model is based on a two-stage network. Motion cues are obtained with…
A parallel dispersive finite-difference time-domain (FDTD) method for the modeling of three-dimensional (3-D) electromagnetic cloaking structures is presented in this paper. The permittivity and permeability of the cloak are mapped to the…
Speech separation remains an important topic for multi-speaker technology researchers. Convolution augmented transformers (conformers) have performed well for many speech processing tasks but have been under-researched for speech…
The articulatory geometric configurations of the vocal tract and the acoustic properties of the resultant speech sound are considered to have a strong causal relationship. This paper aims at finding a joint latent representation between the…
In this work, we present a numerical method that remedies the instabilities of the conventional FDTD approach for solving Maxwell's equations in a space-time dependent magneto-electric medium with direct application to the simulation of the…