中文
相关论文

相关论文: SpatialNet: Extensively Learning Spatial Informati…

200 篇论文

In this paper, we propose a method combining variational autoencoder model of speech with a spatial clustering approach for multi-channel speech separation. The advantage of integrating spatial clustering with a spectral model was shown in…

音频与语音处理 · 电气工程与系统科学 2020-11-25 Katerina Zmolikova , Marc Delcroix , Lukáš Burget , Tomohiro Nakatani , Jan "Honza" Černocký

We present a CNN architecture for speech enhancement from multichannel first-order Ambisonics mixtures. The data-dependent spatial filters, deduced from a mask-based approach, are used to help an automatic speech recognition engine to face…

音频与语音处理 · 电气工程与系统科学 2020-06-03 Amélie Bosca , Alexandre Guérin , Lauréline Perotin , Srđan Kitić

A person tends to generate dynamic attention towards speech under complicated environments. Based on this phenomenon, we propose a framework combining dynamic attention and recursive learning together for monaural speech enhancement. Apart…

声音 · 计算机科学 2020-04-02 Andong Li , Chengshi Zheng , Cunhang Fan , Renhua Peng , Xiaodong Li

Dual-path processing along the temporal and spectral dimensions has shown to be effective in various speech processing applications. While the sound source localization (SSL) models utilizing dual-path processing such as the FN-SSL and…

音频与语音处理 · 电气工程与系统科学 2025-09-24 Yuseon Choi , Hyeonseung Kim , Jewoo Jun , Jong Won Shin

The ambiguity at the boundaries of different semantic classes in point cloud semantic segmentation often leads to incorrect decisions in intelligent perception systems, such as autonomous driving. Hence, accurate delineation of the…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Jiale Chen , Fei Xia , Jianliang Mao , Haoping Wang , Chuanlin Zhang

Neural multi-channel speech enhancement models, in particular those based on the U-Net architecture, demonstrate promising performance and generalization potential. These models typically encode input channels independently, and integrate…

声音 · 计算机科学 2024-10-08 Ibrahim Aldarmaki , Thamar Solorio , Bhiksha Raj , Hanan Aldarmaki

We propose a spatio-spectral, combined model-based and data-driven diarization pipeline consisting of TDOA-based segmentation followed by embedding-based clustering. The proposed system requires neither access to multi-channel training data…

音频与语音处理 · 电气工程与系统科学 2025-09-01 Tobias Cord-Landwehr , Tobias Gburrek , Marc Deegen , Reinhold Haeb-Umbach

We propose a method using a long short-term memory (LSTM) network to estimate the noise power spectral density (PSD) of single-channel audio signals represented in the short time Fourier transform (STFT) domain. An LSTM network common to…

信号处理 · 电气工程与系统科学 2020-11-11 Xiaofei Li , Simon Leglaive , Laurent Girin , Radu Horaud

We propose multi-microphone complex spectral mapping, a simple way of applying deep learning for time-varying non-linear beamforming, for speaker separation in reverberant conditions. We aim at both speaker separation and dereverberation.…

声音 · 计算机科学 2021-05-25 Zhong-Qiu Wang , Peidong Wang , DeLiang Wang

In this paper, we theoretically investigate a new technique for simultaneous information and power transfer (SWIPT) in multiple-input multiple-output (MIMO) point-to-point with radio frequency energy harvesting capabilities. The proposed…

信息论 · 计算机科学 2015-04-10 Stelios Timotheou , Ioannis Krikidis , Sotiris Karachontzitis , Kostas Berberidis

Recent progress in speech separation has been largely driven by advances in deep neural networks, yet their high computational and memory requirements hinder deployment on resource-constrained devices. A significant inefficiency in…

音频与语音处理 · 电气工程与系统科学 2025-07-09 Mohamed Elminshawi , Srikanth Raj Chetupalli , Emanuël A. P. Habets

Recent advancements in multi-scale architectures have demonstrated exceptional performance in image denoising tasks. However, existing architectures mainly depends on a fixed single-input single-output Unet architecture, ignoring the…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Xu Zhao , Chen Zhao , Xiantao Hu , Hongliang Zhang , Ying Tai , Jian Yang

Self-supervised video denoising aims to remove noise from videos without relying on ground truth data, leveraging the video itself to recover clean frames. Existing methods often rely on simplistic feature stacking or apply optical flow…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Zikang Chen , Tao Jiang , Xiaowan Hu , Wang Zhang , Huaqiu Li , Haoqian Wang

Advancements in deep learning and voice-activated technologies have driven the development of human-vehicle interaction. Distributed microphone arrays are widely used in in-car scenarios because they can accurately capture the voices of…

音频与语音处理 · 电气工程与系统科学 2024-09-16 Ziqian Wang , Jiayao Sun , Zihan Zhang , Xingchen Li , Jie Liu , Lei Xie

Most speech separation methods, trying to separate all channel sources simultaneously, are still far from having enough general- ization capabilities for real scenarios where the number of input sounds is usually uncertain and even dynamic.…

声音 · 计算机科学 2021-02-09 Chenxing Li , Jiaming Xu , Nima Mesgarani , Bo Xu

Hyperspectral images (HSIs) have been widely used in a variety of applications thanks to the rich spectral information they are able to provide. Among all HSI processing tasks, HSI denoising is a crucial step. Recently, deep learning-based…

图像与视频处理 · 电气工程与系统科学 2022-02-16 Zhiqiang Wang , Zhenfeng Shao , Xiao Huang , Jiaming Wang , Tao Lu , Sihang Zhang

We propose a new end-to-end neural acoustic model for automatic speech recognition. The model is composed of multiple blocks with residual connections between them. Each block consists of one or more modules with 1D time-channel separable…

音频与语音处理 · 电气工程与系统科学 2019-10-24 Samuel Kriman , Stanislav Beliaev , Boris Ginsburg , Jocelyn Huang , Oleksii Kuchaiev , Vitaly Lavrukhin , Ryan Leary , Jason Li , Yang Zhang

This paper proposes a full-band and sub-band fusion model, named as FullSubNet, for single-channel real-time speech enhancement. Full-band and sub-band refer to the models that input full-band and sub-band noisy spectral feature, output…

音频与语音处理 · 电气工程与系统科学 2024-07-04 Xiang Hao , Xiangdong Su , Radu Horaud , Xiaofei Li

For most of the state-of-the-art speech enhancement techniques, a spectrogram is usually preferred than the respective time-domain raw data since it reveals more compact presentation together with conspicuous temporal information over a…

声音 · 计算机科学 2016-08-24 Syu-Siang Wang , Alan Chern , Yu Tsao , Jeih-weih Hung , Xugang Lu , Ying-Hui Lai , Borching Su

We propose a novel framework for representing neural fields on triangle meshes that is multi-resolution across both spatial and frequency domains. Inspired by the Neural Fourier Filter Bank (NFFB), our architecture decomposes the spatial…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Avigail Cohen Rimon , Tal Shnitzer , Mirela Ben Chen
‹ 上一页 1 8 9 10 下一页 ›