Related papers: Relative Transfer Function Vector Estimation for A…
Reliable fundamental frequency (F 0) and voicing estimation is essential for neural synthesis, yet many pitch extractors depend on large labeled corpora and degrade under realistic recording artifacts. We propose a lightweight, fully…
The mainstream neural text-to-speech(TTS) pipeline is a cascade system, including an acoustic model(AM) that predicts acoustic feature from the input transcript and a vocoder that generates waveform according to the given acoustic feature.…
Audio-visual speech separation methods aim to integrate different modalities to generate high-quality separated speech, thereby enhancing the performance of downstream tasks such as speech recognition. Most existing state-of-the-art (SOTA)…
In Photoacoustic imaging (PAI), the most prevalent beamforming algorithm is delay-and-sum (DAS) due to its simple implementation. However, it results in a low quality image affected by the high level of sidelobes. Coherence factor (CF) can…
The microwave cavity perturbation method is often used to determine material parameters (electric permittivity and magnetic permeability) at high frequencies and it relies on measurement of the resonator parameters. We present a method to…
This paper considers black- and grey-box continuous-time transfer function estimation from frequency response measurements. The first contribution is a bilinear mapping of the original problem from the imaginary axis onto the unitdisk. This…
This study presents an improved quantum teleportation protocol designed to enhance fidelity in noisy environments by combining weak measurements (WMs) with flip and reversal operations. In our scheme, Alice prepares a four-qubit entangled…
The paper presents a method for audio-based vehicle counting (VC) in low-to-moderate traffic using one-channel sound. We formulate VC as a regression problem, i.e., we predict the distance between a vehicle and the microphone. Minima of the…
A method of interpolating the acoustic transfer function (ATF) between regions that takes into account both the physical properties of the ATF and the directionality of region configurations is proposed. Most spatial ATF interpolation…
In recent years, there has been a growing interest in designing small-footprint yet effective Connectionist Temporal Classification based keyword spotting (CTC-KWS) systems. They are typically deployed on low-resource computing platforms,…
The Continuous Wavelet Transform (CWT) is an effective tool for feature extraction in acoustic recognition using Convolutional Neural Networks (CNNs), particularly when applied to non-stationary audio. However, its high computational cost…
Speech separation models are used for isolating individual speakers in many speech processing applications. Deep learning models have been shown to lead to state-of-the-art (SOTA) results on a number of speech separation benchmarks. One…
Currently, the sub-60 Hz sensitivity of gravitational-wave (GW) detectors like Advanced LIGO is limited by the control noises from auxiliary degrees of freedom, which nonlinearly couple to the main GW readout. One particularly promising way…
In this paper we address the problem of enhancing speech signals in noisy mixtures using a source separation approach. We explore the use of neural networks as an alternative to a popular speech variance model based on supervised…
Numerically, a theoretical analysis of the noise impact caused by spontaneous Raman scattering, four-wave mixing, and linear channel crosstalk on the measurement-device-independent continuous variable quantum key distribution systems is…
There are two types of methods for non-autoregressive text-to-speech models to learn the one-to-many relationship between text and speech effectively. The first one is to use an advanced generative framework such as normalizing flow (NF).…
Convolutional Neural Networks (CNNs) can learn effective features, though have been shown to suffer from a performance drop when the distribution of the data changes from training to test data. In this paper we analyze the internal…
We consider the problem of estimating the covariance matrix of a random signal observed through unknown translations (modeled by cyclic shifts) and corrupted by noise. Solving this problem allows to discover low-rank structures masked by…
Most speech enhancement algorithms make use of the short-time Fourier transform (STFT), which is a simple and flexible time-frequency decomposition that estimates the short-time spectrum of a signal. However, the duration of short STFT…
The imperfections of a receiver's detector affect the performance of two-way continuous-variable quantum key distribution protocols and are difficult to adjust in practical situations. We propose a method to improve the performance of…