Related papers: A comparative study of two-dimensional vocal tract…
In this paper, a semi-discrete spatial finite volume (FV) method is proposed and analyzed for approximating solutions of anomalous subdiffusion equations involving a temporal fractional derivative of order $\alpha \in (0,1)$ in a…
We introduce a dual-wavelength Fourier ptychographic topography (FPT) method that extends the lambda/2 height-range limit of single-wavelength FPT. By reconstructing complex fields at two illumination wavelengths and exploiting their phase…
2.5D cartoon models are methods to simulate three-dimensional (3D)-like movements, such as out-of-plane rotation, from two-dimensional (2D) shapes in different views. However, cartoon objects and characters have several distorted parts…
A finite-difference time-domain (FDTD) modelling of finite-size zero thickness space-time modulated Huygens' metasurfaces based on Generalized Sheet Transition Conditions (GSTCs), is proposed and numerically demonstrated. A typical…
Recent advancements in the field of Diffusion Transformers have substantially improved the generation of high-quality 2D images, 3D videos, and 3D shapes. However, the effectiveness of the Transformer architecture in the domain of co-speech…
The accuracy of finite-difference time-domain (FDTD) modelling of left-handed metamaterials (LHMs) is dramatically improved by using an averaging technique along the boundaries of LHM slabs. The material frequency dispersion of LHMs is…
Recently, frequency domain all-neural beamforming methods have achieved remarkable progress for multichannel speech separation. In parallel, the integration of time domain network structure and beamforming also gains significant attention.…
The mainstream neural text-to-speech(TTS) pipeline is a cascade system, including an acoustic model(AM) that predicts acoustic feature from the input transcript and a vocoder that generates waveform according to the given acoustic feature.…
Low-dose computed tomography (LDCT) reduces radiation exposure but suffers from image artifacts and loss of detail due to quantum and electronic noise, potentially impacting diagnostic accuracy. Transformer combined with diffusion models…
Deep neural networks have been applied to audio spectrograms for respiratory sound classification. Existing models often treat the spectrogram as a synthetic image while overlooking its physical characteristics. In this paper, a Multi-View…
In volume-to-volume translations in medical images, existing models often struggle to capture the inherent volumetric distribution using 3D voxelspace representations, due to high computational dataset demands. We present Score-Fusion, a…
Real-time Magnetic Resonance Imaging (rtMRI) visualizes vocal tract action, offering a comprehensive window into speech articulation. However, its signals are high dimensional and noisy, hindering interpretation. We investigate compact…
${\tt simwave}$ is an open-source Python package to perform wave simulations in 2D or 3D domains. It solves the constant and variable density acoustic wave equation with the finite difference method and has support for domain truncation…
Speech enhancement in multichannel settings has been realized by utilizing the spatial information embedded in multiple microphone signals. Moreover, deep neural networks (DNNs) have been recently advanced in this field; however, studies on…
In this paper, a spatially dispersive finite-difference time-domain (FDTD) method to model wire media is developed and validated. Sub-wavelength imaging properties of the finite wire medium slabs are examined. It is demonstrated that the…
With the advances in deep learning, the performance of end-to-end (E2E) single-task models for speech and audio processing has been constantly improving. However, it is still challenging to build a general-purpose model with high…
Recent advancements in diffusion probabilistic models (DPMs) have revolutionized image processing, demonstrating significant potential in medical applications. Accurate segmentation of the left ventricle (LV) in echocardiograms is crucial…
A novel 3-D higher-order finite-difference time-domain framework with complex frequency-shifted perfectly matched layer for the modeling of wave propagation in cold plasma is presented. Second- and fourth-order spatial approximations are…
From nano-scale heat transfer point of view, currently one of the most interesting and challenging tasks is to quantitatively analyzing phonon mode specific transport properties in solid materials, which plays vital role in many emerging…
There has been a growing interest in using end-to-end acoustic models for singing voice synthesis (SVS). Typically, these models require an additional vocoder to transform the generated acoustic features into the final waveform. However,…