English
Related papers

Related papers: Time-Varying Quasi-Closed-Phase Analysis for Accur…

200 papers

Formant tracking is investigated in this study by using trackers based on dynamic programming (DP) and deep neural nets (DNNs). Using the DP approach, six formant estimation methods were first compared. The six methods include linear…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-06 Dhananjaya Gowda , Bajibabu Bollepalli , Sudarsana Reddy Kadiri , Paavo Alku

In this study, formant tracking is investigated by refining the formants tracked by an existing data-driven tracker, DeepFormants, using the formants estimated in a model-driven manner by linear prediction (LP)-based methods. As LP-based…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-21 Paavo Alku , Sudarsana Reddy Kadiri , Dhananjaya Gowda

Deep learning has brought significant improvements to the field of cross-modal representation learning. For tasks such as text-to-speech (TTS), voice conversion (VC), and automatic speech recognition (ASR), a cross-modal fine-grained…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-29 Chunyu Qiang , Wang Geng , Yi Zhao , Ruibo Fu , Tao Wang , Cheng Gong , Tianrui Wang , Qiuyu Liu , Jiangyan Yi , Zhengqi Wen , Chen Zhang , Hao Che , Longbiao Wang , Jianwu Dang , Jianhua Tao

Formants are the spectral maxima that result from acoustic resonances of the human vocal tract, and their accurate estimation is among the most fundamental speech processing problems. Recent work has been shown that those frequencies can…

Sound · Computer Science 2022-06-24 Yosi Shrem , Felix Kreuk , Joseph Keshet

Vocal tract resonance characteristics in acoustic speech signals are classically tracked using frame-by-frame point estimates of formant frequencies followed by candidate selection and smoothing using dynamic programming methods that…

Applications · Statistics 2012-10-15 Daryush D. Mehta , Daniel Rudoy , Patrick J. Wolfe

We study signal processing tasks in which the signal is mapped via some generalized time-frequency transform to a higher dimensional time-frequency space, processed there, and synthesized to an output signal. We show how to approximate such…

Numerical Analysis · Mathematics 2021-09-07 Ron Levie , Haim Avron , Gitta Kutyniok

The phase vocoder (PV) is a widely spread technique for processing audio signals. It employs a short-time Fourier transform (STFT) analysis-modify-synthesis loop and is typically used for time-scaling of signals by means of using different…

Sound · Computer Science 2022-02-16 Zdenek Prusa , Nicki Holighaus

This paper presents a novel adaptive fading cubature Kalman filter (AFCKF) based on double transitive factors. The developed adaptive algorithm is explained in two stages; stage (i) a single transitive factor is used to update the predicted…

Systems and Control · Electrical Eng. & Systems 2021-08-26 Mundla Narasimhappa

We present a method for blind acoustic parameter estimation from single-channel reverberant speech. The method is structured into three stages. In the first stage, a variational auto-encoder is trained to extract latent representations of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-30 Philipp Götz , Cagdas Tuna , Andreas Brendel , Andreas Walther , Emanuël A. P. Habets

In recent years, deep learning-based approaches have significantly improved the performance of single-channel speech enhancement. However, due to the limitation of training data and computational complexity, real-time enhancement of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-16 Zehua Zhang , Lu Zhang , Xuyi Zhuang , Yukun Qian , Heng Li , Mingjiang Wang

We present a single-channel phase-sensitive speech enhancement algorithm that is based on modulation-domain Kalman filtering and on tracking the speech phase using circular statistics. With Kalman filtering, using that speech and noise are…

Sound · Computer Science 2017-08-08 Nikolaos Dionelis , Mike Brookes

This paper presents two methods for approximating the performance of coded multicarrier systems operating over frequency-selective, quasi-static fading channels with non-ideal interleaving. The first method is based on approximating the…

Information Theory · Computer Science 2007-07-13 C. Snow , L. Lampe , R. Schober

Non-conventional receivers for phase-coherent states based on non-Gaussian measurements such as photon counting surpass the sensitivity limits of shot-noise-limited coherent receivers, the quantum noise limit (QNL). These non-Gaussian…

Quantum Physics · Physics 2020-07-07 M. T. DiMario , F. E. Becerra

We consider the problem of quantitative predictive monitoring (QPM) of stochastic systems, i.e., predicting at runtime the degree of satisfaction of a desired temporal logic property from the current state of the system. Since computational…

Artificial Intelligence · Computer Science 2025-09-03 Francesca Cairoli , Luca Bortolussi , Jyotirmoy V. Deshmukh , Lars Lindemann , Nicola Paoletti

Vocoders, encoding speech signals into acoustic features and allowing for speech signal reconstruction from them, have been studied for decades. Recently, the rise of deep learning has particularly driven the development of neural vocoders…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-03 Shaowen Chen , Tomoki Toda

To reduce the influence of random channel polarization variation, especially fast polarization perturbation,for continuous-variable quantum key distribution (CV-QKD) systems, a simple and fast polarization tracking algorithm is proposed and…

Quantum Physics · Physics 2023-03-22 Yan Pan , Heng Wang , Yun Shao , Yaodi Pi , Ting Ye , Shuai Zhang , Yang Li , Wei Huang , Bingjie Xu

Partial deepfake speech detection requires identifying manipulated regions that may occur within short temporal portions of an otherwise bona fide utterance, making the task particularly challenging for conventional utterance-level…

Sound · Computer Science 2026-04-06 Inbal Rimon , Oren Gal , Haim Permuter

Tracking cells in time-lapse videos is an essential technique for monitoring cell population dynamics at a single-cell level. Current methods for cell tracking are developed on videos with mostly single, constant signals and do not detect…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Florian Bürger , Martim Dias Gomes , Nica Gutu , Adrián E. Granada , Noémie Moreau , Katarzyna Bozek

Complex-valued processing has brought deep learning-based speech enhancement and signal extraction to a new level. Typically, the process is based on a time-frequency (TF) mask which is applied to a noisy spectrogram, while complex masks…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-02 Hendrik Schröter , Alberto N. Escalante-B. , Tobias Rosenkranz , Andreas Maier

We propose a novel neural waveform compression method to catalyze emerging speech semantic communications. By introducing nonlinear transform and variational modeling, we effectively capture the dependencies within speech frames and…

Sound · Computer Science 2022-12-14 Shengshi Yao , Zixuan Xiao , Sixian Wang , Jincheng Dai , Kai Niu , Ping Zhang
‹ Prev 1 2 3 10 Next ›