中文
相关论文

相关论文: Deep Filter Estimation from Inter-Frame Correlatio…

200 篇论文

Several speech processing systems have demonstrated considerable performance improvements when deep complex neural networks (DCNN) are coupled with self-attention (SA) networks. However, the majority of DCNN-based studies on speech…

音频与语音处理 · 电气工程与系统科学 2022-11-24 Vinay Kothapally , John H. L. Hansen

Robust audio anti-spoofing has been increasingly challenging due to the recent advancements on deepfake techniques. While spectrograms have demonstrated their capability for anti-spoofing, complementary information presented in multi-order…

声音 · 计算机科学 2024-10-04 Penghui Wen , Kun Hu , Wenxi Yue , Sen Zhang , Wanlei Zhou , Zhiyong Wang

In recent years, deep networks have led to dramatic improvements in speech enhancement by framing it as a data-driven pattern recognition problem. In many modern enhancement systems, large amounts of data are used to train a deep network to…

How to effectively explore spatial-temporal features is important for video colorization. Instead of stacking multiple frames along the temporal dimension or recurrently propagating estimated features that will accumulate errors or cannot…

计算机视觉与模式识别 · 计算机科学 2024-04-10 Yixin Yang , Jiangxin Dong , Jinhui Tang , Jinshan Pan

Direct-path relative transfer function (DP-RTF) refers to the ratio between the direct-path acoustic transfer functions of two microphone channels. Though DP-RTF fully encodes the sound spatial cues and serves as a reliable localization…

声音 · 计算机科学 2022-02-17 Bing Yang , Hong Liu , Xiaofei Li

Time of Flight (ToF) is a prevalent depth sensing technology in the fields of robotics, medical imaging, and non-destructive testing. Yet, ToF sensing faces challenges from complex ambient conditions making an inverse modelling from the…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Christopher Hahne , Michel Hayoz , Raphael Sznitman

This paper presents EffortNet, a novel deep learning framework for decoding individual listening effort from electroencephalography (EEG) during speech comprehension. Listening effort represents a significant challenge in speech-hearing…

音频与语音处理 · 电气工程与系统科学 2025-08-22 Ching-Chih Sung , Cheng-Hung Hsin , Yu-Anne Shiah , Bo-Jyun Lin , Yi-Xuan Lai , Chia-Ying Lee , Yu-Te Wang , Borchin Su , Yu Tsao

Audio-visual speech separation aims to isolate each speaker's clean voice from mixtures by leveraging visual cues such as lip movements and facial features. While visual information provides complementary semantic guidance, existing methods…

声音 · 计算机科学 2025-10-13 Ke Xue , Rongfei Fan , Lixin , Dawei Zhao , Chao Zhu , Han Hu

Extracting direct-path spatial feature is crucial for sound source localization in adverse acoustic environments. This paper proposes the IPDnet, a neural network that estimates direct-path inter-channel phase difference (DP-IPD) of sound…

音频与语音处理 · 电气工程与系统科学 2024-05-14 Yabo Wang , Bing Yang , Xiaofei Li

Deep learning algorithm are increasingly used for speech enhancement (SE). In supervised methods, global and local information is required for accurate spectral mapping. A key restriction is often poor capture of key contextual information.…

声音 · 计算机科学 2022-10-28 Jianqiao Cui , Stefan Bleeck

Real-world audio recordings are often degraded by factors such as noise, reverberation, and equalization distortion. This paper introduces HiFi-GAN, a deep learning method to transform recorded speech to sound as though it had been recorded…

音频与语音处理 · 电气工程与系统科学 2020-09-23 Jiaqi Su , Zeyu Jin , Adam Finkelstein

Speaker verification systems have seen significant advancements with the introduction of Multi-scale Feature Aggregation (MFA) architectures, such as MFA-Conformer and ECAPA-TDNN. These models leverage information from various network…

声音 · 计算机科学 2024-10-08 Satvik Dixit , Massa Baali , Rita Singh , Bhiksha Raj

We demonstrate a first example for employing deep learning in predicting frame errors for a Collaborative Intelligent Radio Network (CIRN) using a dataset collected during participation in the final scrimmages of the DARPA SC2 challenge.…

信号处理 · 电气工程与系统科学 2020-12-29 Abu Shafin Mohammad Mahdee Jameel , Ahmed P. Mohamed , Xiwen Zhang , Aly El Gamal

Real-image super-resolution (Real-ISR) seeks to recover HR images from LR inputs with mixed, unknown degradations. While diffusion models surpass GANs in perceptual quality, they under-reconstruct high-frequency (HF) details due to a…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Seungho Choi , Jeahun Sung , Jihyong Oh

Echo path delay (or ref-delay) estimation is a big challenge in acoustic echo cancellation. Different devices may introduce various ref-delay in practice. Ref-delay inconsistency slows down the convergence of adaptive filters, and also…

音频与语音处理 · 电气工程与系统科学 2022-08-12 Yi Zhang , Chengyun Deng , Shiqian Ma , Yongtao Sha , Hui Song

A novel framework, called InterGridNet, is introduced, leveraging a shallow RawNet model for geolocation classification of Electric Network Frequency (ENF) signatures in the SP Cup 2016 dataset. During data preparation, recordings are…

In speaker verification systems, the utilization of short utterances presents a persistent challenge, leading to performance degradation primarily due to insufficient phonetic information to characterize the speakers. To overcome this…

音频与语音处理 · 电气工程与系统科学 2024-06-12 Seung-bin Kim , Chan-yeong Lim , Jungwoo Heo , Ju-ho Kim , Hyun-seo Shin , Kyo-Won Koo , Ha-Jin Yu

Multi-channel speech enhancement with ad-hoc sensors has been a challenging task. Speech model guided beamforming algorithms are able to recover natural sounding speech, but the speech models tend to be oversimplified or the inference would…

计算与语言 · 计算机科学 2018-02-16 Kaizhi Qian , Yang Zhang , Shiyu Chang , Xuesong Yang , Dinei Florencio , Mark Hasegawa-Johnson

Image deblurring plays a crucial role in enhancing visual clarity across various applications. Although most deep learning approaches primarily focus on sRGB images, which inherently lose critical information during the image signal…

图像与视频处理 · 电气工程与系统科学 2025-09-22 Wenlong Jiao , Binglong Li , Wei Shang , Ping Wang , Dongwei Ren

Single-channel speech dereverberation aims at extracting a dry speech signal from a recording affected by the acoustic reflections in a room. However, most current deep learning-based approaches for speech dereverberation are not…

声音 · 计算机科学 2024-07-12 Louis Bahrman , Mathieu Fontaine , Jonathan Le Roux , Gaël Richard