中文
相关论文

相关论文: On the Parameter Estimation of Sinusoidal Models f…

200 篇论文

We propose an end-to-end model based on convolutional and recurrent neural networks for speech enhancement. Our model is purely data-driven and does not make any assumptions about the type or the stationarity of the noise. In contrast to…

声音 · 计算机科学 2018-05-03 Han Zhao , Shuayb Zarar , Ivan Tashev , Chin-Hui Lee

Integrating front-end speech enhancement (SE) models with self-supervised learning (SSL)-based speech models is effective for downstream tasks in noisy conditions. SE models are commonly fine-tuned using SSL representations with mean…

计算与语言 · 计算机科学 2026-01-30 Amit Meghanani , Thomas Hain

Form about four decades human beings have been dreaming of an intelligent machine which can master the natural speech. In its simplest form, this machine should consist of two subsystems, namely automatic speech recognition (ASR) and speech…

声音 · 计算机科学 2013-05-08 Urmila Shrawankar , V. M. Thakare

Subword units are commonly used for end-to-end automatic speech recognition (ASR), while a fully acoustic-oriented subword modeling approach is somewhat missing. We propose an acoustic data-driven subword modeling (ADSM) approach that…

计算与语言 · 计算机科学 2023-10-24 Wei Zhou , Mohammad Zeineldeen , Zuoyun Zheng , Ralf Schlüter , Hermann Ney

We present a novel method for extracting neural embeddings that model the background acoustics of a speech signal. The extracted embeddings are used to estimate specific parameters related to the background acoustic properties of the signal…

音频与语音处理 · 电气工程与系统科学 2024-06-11 Sri Harsha Dumpala , Dushyant Sharma , Chandramouli Shama Sastri , Stanislav Kruchinin , James Fosburgh , Patrick A. Naylor

Speech enhancement (SE) aims to extract the clean waveform from noise-contaminated measurements to improve the speech quality and intelligibility. Although learning-based methods can perform much better than traditional counterparts, the…

音频与语音处理 · 电气工程与系统科学 2024-09-23 Haoyin Yan , Jie Zhang , Cunhang Fan , Yeping Zhou , Peiqi Liu

Various parametric representations have been proposed to model the speech signal. While the performance of such vocoders is well-known in the context of speech processing, their extrapolation to singing voice synthesis might not be…

音频与语音处理 · 电气工程与系统科学 2020-06-09 Onur Babacan , Thomas Drugman , Tuomo Raitio , Daniel Erro , Thierry Dutoit

One of the biggest challenges in multi-microphone applications is the estimation of the parameters of the signal model such as the power spectral densities (PSDs) of the sources, the early (relative) acoustic transfer functions of the…

音频与语音处理 · 电气工程与系统科学 2018-10-16 Andreas I. Koutrouvelis , Richard C. Hendriks , Richard Heusdens , Jesper Jensen

In this paper, we attempt to study the conditioning of the Spherical Harmonic Matrix (SHM), which is widely used in the discrete, limited order orthogonal representation of sound fields. SHM's has been widely used in the audio applications…

音频与语音处理 · 电气工程与系统科学 2018-03-07 C Sandeep Reddy , Rajesh M Hegde

Synthetically generated speech has rapidly approached human levels of naturalness. However, the paradox remains that ASR systems, when trained on TTS output that is judged as natural by humans, continue to perform badly on real speech. In…

音频与语音处理 · 电气工程与系统科学 2024-10-17 Christoph Minixhofer , Ondrej Klejch , Peter Bell

Recent synthetic speech detection models typically adapt a pre-trained SSL model via finetuning, which is computationally demanding. Parameter-Efficient Fine-Tuning (PEFT) offers an alternative. However, existing methods lack the specific…

声音 · 计算机科学 2025-10-30 Yassine El Kheir , Fabian Ritter-Guttierez , Arnab Das , Tim Polzehl , Sebastian Möller

End-to-end Speech-to-text Translation (E2E-ST), which directly translates source language speech to target language text, is widely useful in practice, but traditional cascaded approaches (ASR+MT) often suffer from error propagation in the…

计算与语言 · 计算机科学 2021-02-10 Junkun Chen , Mingbo Ma , Renjie Zheng , Liang Huang

We consider signals that follow a parametric distribution where the parameter values are unknown. To estimate such signals from noisy measurements in scalar channels, we study the empirical performance of an empirical Bayes (EB) approach…

信息论 · 计算机科学 2014-05-12 Yanting Ma , Jin Tan , Nikhil Krishnan , Dror Baron

We present the first application of the extended Fast Action Minimization method (eFAM) to a real dataset, the SDSS-DR12 Combined Sample, to reconstruct galaxies orbits back-in-time, their two-point correlation function (2PCF) in…

宇宙学与河外天体物理 · 物理学 2021-02-17 E. Sarpa , A. Veropalumbo , C. Schimd , E. Branchini , S. Matarrese

The field of deep-learning-based ECG analysis has been largely dominated by convolutional architectures. This work explores the prospects of applying the recently introduced structured state space models (SSMs) as a particularly promising…

机器学习 · 计算机科学 2022-11-15 Temesgen Mehari , Nils Strodthoff

Far-field speech recognition in noisy and reverberant conditions remains a challenging problem despite recent deep learning breakthroughs. This problem is commonly addressed by acquiring a speech signal from multiple microphones and…

音频与语音处理 · 电气工程与系统科学 2018-10-17 Zhong Meng , Shinji Watanabe , John R. Hershey , Hakan Erdogan

Spatial consistency was proposed in the 3GPP TR 38.901 channel model to ensure that closely spaced mobile terminals have similar channels. Future extensions of this model might incorporate mobility at both ends of the link. This requires…

信号处理 · 电气工程与系统科学 2019-07-01 Stephan Jaeckel , Leszek Raschkowski , Frank Burkhardt , Lars Thiele

In this paper, we propose to utilise diffusion models for data augmentation in speech emotion recognition (SER). In particular, we present an effective approach to utilise improved denoising diffusion probabilistic models (IDDPM) to…

声音 · 计算机科学 2023-05-22 Ibrahim Malik , Siddique Latif , Raja Jurdak , Björn Schuller

We develop a Bayesian framework for sensing which adapts the sensing time and/or basis functions to the instantaneous sensing quality measured in terms of the expected posterior mean-squared error. For sparse Gaussian sources a significant…

信息论 · 计算机科学 2018-02-12 Ralf R. Müller , Ali Bereyhi , Christoph F. Mecklenbräuker

We consider estimating an unknown signal, both blocky and sparse, which is corrupted by additive noise. We study three interrelated least squares procedures and their asymptotic properties. The first procedure is the fused lasso, put…

统计理论 · 数学 2009-08-31 Alessandro Rinaldo