相关论文: An extended two-dimensional vocal tract model for …
We introduce SoundSpaces 2.0, a platform for on-the-fly geometry-based audio rendering for 3D environments. Given a 3D mesh of a real-world environment, SoundSpaces can generate highly realistic acoustics for arbitrary sounds captured from…
A time-domain numerical modeling of transversely isotropic Biot poroelastic waves is proposed in two dimensions. The viscous dissipation occurring in the pores is described using the dynamic permeability model developed by…
With the advances in deep learning, the performance of end-to-end (E2E) single-task models for speech and audio processing has been constantly improving. However, it is still challenging to build a general-purpose model with high…
The articulatory geometric configurations of the vocal tract and the acoustic properties of the resultant speech sound are considered to have a strong causal relationship. This paper aims at finding a joint latent representation between the…
To enhance the performance of end-to-end (E2E) speech recognition systems in noisy or low signal-to-noise ratio (SNR) conditions, this paper introduces NoisyD-CT, a novel tri-stage training framework built on the Conformer-Transducer…
Time-distance helioseismology is the method of the study of the propagation of waves through the solar interior via the travel times of those waves. The travel times of wave packets contain information about the conditions in the interior…
Recent high-performance transformer-based speech enhancement models demonstrate that time domain methods could achieve similar performance as time-frequency domain methods. However, time-domain speech enhancement systems typically receive…
Using analytical expressions for the pressure and velocity waveforms in tapered vessels, we construct a linear 1D model for wave propagation in stenotic vessels in the frequency domain. We demonstrate that using only two parameters to…
Diffusion models have recently emerged as a powerful technique in image generation, especially for image super-resolution tasks. While 2D diffusion models significantly enhance the resolution of individual images, existing diffusion-based…
Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hindered their applications to speech synthesis. This paper…
Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hinder their applications to text-to-speech deployment. Through…
Speech enhancement is designed to enhance the intelligibility and quality of speech across diverse noise conditions. Recently, diffusion model has gained lots of attention in speech enhancement area, achieving competitive results. Current…
At millimeter-wave frequencies, diffuse scattering from rough surfaces is an important propagation mechanism. Including this mechanism in radio propagation modeling tools,such as ray-tracing, is a key step towards realizing accurate…
Robust voice activity detection (VAD) is a challenging task in low signal-to-noise (SNR) environments. Recent studies show that speech enhancement is helpful to VAD, but the performance improvement is limited. To address this issue, here we…
Two-stage pipeline is popular in speech enhancement tasks due to its superiority over traditional single-stage methods. The current two-stage approaches usually enhance the magnitude spectrum in the first stage, and further modify the…
The sound-localization and, in particular, biosonar system of toothed whales is exceptionally performant. How this is achieved is not clear, given that: (i) toothed whales have no pinnae; (ii) while their auditory pathways have been studied…
In the current state of 3D object detection research, the severe scarcity of annotated 3D data, substantial disparities across different data modalities, and the absence of a unified architecture, have impeded the progress towards the goal…
The vector representations of fixed dimensionality for words (in text) offered by Word2Vec have been shown to be very useful in many application scenarios, in particular due to the semantic information they carry. This paper proposes a…
Speech separation remains an important topic for multi-speaker technology researchers. Convolution augmented transformers (conformers) have performed well for many speech processing tasks but have been under-researched for speech…
We develop a weakly nonlinear model of duct acoustics in two and three dimensions (without flow). The work extends the previous work of McTavish & Brambley (2019, J. Fluid Mech. 875, pp. 411-447) to three dimensions and significantly…