中文
相关论文

相关论文: Long-frame-shift Neural Speech Phase Prediction wi…

200 篇论文

Recently, progressive learning has shown its capacity to improve speech quality and speech intelligibility when it is combined with deep neural network (DNN) and long short-term memory (LSTM) based monaural speech enhancement algorithms,…

声音 · 计算机科学 2020-01-14 Andong Li , Minmin Yuan , Chengshi Zheng , Xiaodong Li

Speech enhancement techniques based on deep learning have brought significant improvement on speech quality and intelligibility. Nevertheless, a large gain in speech quality measured by objective metrics, such as perceptual evaluation of…

音频与语音处理 · 电气工程与系统科学 2020-07-06 Bo Wu , Meng Yu , Lianwu Chen , Yong Xu , Chao Weng , Dan Su , Dong Yu

Recent deep learning approaches have achieved impressive performance on speech enhancement and separation tasks. However, these approaches have not been investigated for separating mixtures of arbitrary sounds of different types, a task we…

Segments that span contiguous parts of inputs, such as phonemes in speech, named entities in sentences, actions in videos, occur frequently in sequence prediction problems. Segmental models, a class of models that explicitly hypothesizes…

计算与语言 · 计算机科学 2018-06-14 Hao Tang

Multi-channel speech enhancement extracts speech using multiple microphones that capture spatial cues. Effectively utilizing directional information is key for multi-channel enhancement. Deep learning shows great potential on multi-channel…

声音 · 计算机科学 2023-09-21 Jiahui Pan , Pengjie Shen , Hui Zhang , Xueliang Zhang

Speech intelligibility can be degraded due to multiple factors, such as noisy environments, technical difficulties or biological conditions. This work is focused on the development of an automatic non-intrusive system for predicting the…

音频与语音处理 · 电气工程与系统科学 2024-02-07 Miguel Fernández-Díaz , Ascensión Gallardo-Antolín

In this paper, we propose a transformer-based architecture, called two-stage transformer neural network (TSTNN) for end-to-end speech denoising in the time domain. The proposed model is composed of an encoder, a two-stage transformer module…

音频与语音处理 · 电气工程与系统科学 2021-03-19 Kai Wang , Bengbeng He , Wei-Ping Zhu

Nonnegative Matrix Factorization (NMF) is a powerful tool for decomposing mixtures of audio signals in the Time-Frequency (TF) domain. In applications such as source separation, the phase recovery for each extracted component is a major…

声音 · 计算机科学 2016-11-17 Paul Magron , Roland Badeau , Bertrand David

Vision-and-language navigation (VLN) is a crucial but challenging cross-modal navigation task. One powerful technique to enhance the generalization performance in VLN is the use of an independent speaker model to provide pseudo instructions…

计算机视觉与模式识别 · 计算机科学 2024-03-07 Liuyi Wang , Chengju Liu , Zongtao He , Shu Li , Qingqing Yan , Huiyi Chen , Qijun Chen

The Short-Time Fourier Transform (STFT) has been a staple of signal processing, often being the first step for many audio tasks. A very familiar process when using the STFT is the search for the best STFT parameters, as they often have…

音频与语音处理 · 电气工程与系统科学 2021-05-17 An Zhao , Krishna Subramani , Paris Smaragdis

This paper addresses the problem of speech separation and enhancement from multichannel convolutive and noisy mixtures, \emph{assuming known mixing filters}. We propose to perform the speech separation and enhancement task in the short-time…

声音 · 计算机科学 2019-01-31 Xiaofei Li , Laurent Girin , Sharon Gannot , Radu Horaud

While deep reinforcement learning (RL) has been demonstrated effective in solving complex control tasks, sample efficiency remains a key challenge due to the large amounts of data required for remarkable performance. Existing research…

机器学习 · 计算机科学 2023-10-25 Mingxuan Ye , Yufei Kuang , Jie Wang , Rui Yang , Wengang Zhou , Houqiang Li , Feng Wu

Speaker localization using microphone arrays depends on accurate time delay estimation techniques. For decades, methods based on the generalized cross correlation with phase transform (GCC-PHAT) have been widely adopted for this purpose.…

音频与语音处理 · 电气工程与系统科学 2022-09-22 Axel Berg , Mark O'Connor , Kalle Åström , Magnus Oskarsson

In multi-temporal SAR interferometry (MT-InSAR), persistent scatterer (PS) pixels are used to estimate geophysical parameters, essentially deformation. Conventionally, PS pixels are selected on the basis of the estimated noise present in…

图像与视频处理 · 电气工程与系统科学 2020-03-12 Ashutosh Tiwari , Avadh Bihari Narayan , Onkar Dikshit

Multi-frame algorithms for single-microphone speech enhancement, e.g., the multi-frame minimum variance distortionless response (MFMVDR) filter, are able to exploit speech correlation across adjacent time frames in the short-time Fourier…

音频与语音处理 · 电气工程与系统科学 2021-05-17 Marvin Tammen , Simon Doclo

Statistical signal processing based speech enhancement methods adopt expert knowledge to design the statistical models and linear filters, which is complementary to the deep neural network (DNN) based methods which are data-driven. In this…

音频与语音处理 · 电气工程与系统科学 2021-04-19 Wei Xue , Gang Quan , Chao Zhang , Guohong Ding , Xiaodong He , Bowen Zhou

This paper addresses the problem of under-determinded speech source separation from multichannel microphone singals, i.e. the convolutive mixtures of multiple sources. The time-domain signals are first transformed to the short-time Fourier…

声音 · 计算机科学 2019-04-11 Xiaofei Li , Laurent Girin , Radu Horaud

Accurate alignment of dysfluent speech with intended text is crucial for automating the diagnosis of neurodegenerative speech disorders. Traditional methods often fail to model phoneme similarities effectively, limiting their performance.…

Neural vocoders are central to speech synthesis; despite their success, most still suffer from limited prosody modeling and inaccurate phase reconstruction. We propose a vocoder that introduces prosody-guided harmonic attention to enhance…

声音 · 计算机科学 2026-01-22 Mohammed Salah Al-Radhi , Riad Larbi , Mátyás Bartalis , Géza Németh

Various informative factors mixed in speech signals, leading to great difficulty when decoding any of the factors. An intuitive idea is to factorize each speech frame into individual informative factors, though it turns out to be highly…

音频与语音处理 · 电气工程与系统科学 2018-03-05 Lantian Li , Dong Wang , Yixiang Chen , Ying Shi , Zhiyuan Tang , Thomas Fang Zheng
‹ 上一页 1 8 9 10 下一页 ›