中文
相关论文

相关论文: TT-Net: Dual-path transformer based sound field tr…

200 篇论文

Sub-band models have achieved promising results due to their ability to model local patterns in the spectrogram. Some studies further improve the performance by fusing sub-band and full-band information. However, the structure for the…

声音 · 计算机科学 2022-01-26 Feng Dang , Hangting Chen , Pengyuan Zhang

Neural beamformers, which integrate both pre-separation and beamforming modules, have demonstrated impressive effectiveness in target speech extraction. Nevertheless, the performance of these beamformers is inherently limited by the…

声音 · 计算机科学 2023-09-08 Aoqi Guo , Sichong Qian , Baoxiang Li , Dazhi Gao

This paper presents SHTNet, a lightweight spherical harmonic transform (SHT) based framework, which is designed to address cross-array generalization challenges in multi-channel automatic speech recognition (ASR) through three key…

音频与语音处理 · 电气工程与系统科学 2025-10-22 Xiangzhu Kong , Huang Hao , Zhijian Ou

Multi-channel speech enhancement utilizes spatial information from multiple microphones to extract the target speech. However, most existing methods do not explicitly model spatial cues, instead relying on implicit learning from…

声音 · 计算机科学 2023-09-20 Jiahui Pan , Shulin He , Hui Zhang , Xueliang Zhang

Most of the current deep learning-based approaches for speech enhancement only operate in the spectrogram or waveform domain. Although a cross-domain transformer combining waveform- and spectrogram-domain inputs has been proposed, its…

声音 · 计算机科学 2023-10-31 Jialu Li , Junhui Li , Pu Wang , Youshan Zhang

Multi-channel speech enhancement extracts speech using multiple microphones that capture spatial cues. Effectively utilizing directional information is key for multi-channel enhancement. Deep learning shows great potential on multi-channel…

声音 · 计算机科学 2023-09-21 Jiahui Pan , Pengjie Shen , Hui Zhang , Xueliang Zhang

For the accurate representation and reconstruction of band-limited signals on the sphere, an optimal-dimensionality sampling scheme has been recently proposed which requires the optimal number of samples equal to the number of degrees of…

信息论 · 计算机科学 2017-09-11 Wajeeha Nafees , Zubair Khalid , Rodney A. Kennedy , Jason D. McEwen

This paper introduces a dual-signal transformation LSTM network (DTLN) for real-time speech enhancement as part of the Deep Noise Suppression Challenge (DNS-Challenge). This approach combines a short-time Fourier transform (STFT) and a…

音频与语音处理 · 电气工程与系统科学 2020-10-23 Nils L. Westhausen , Bernd T. Meyer

Utilizing spherical harmonic (SH) domain has been established as the default method of obtaining continuity over space in head-related transfer functions (HRTFs). This paper concerns different variants of extending this solution by…

音频与语音处理 · 电气工程与系统科学 2023-07-19 Adam Szwajcowski

Deep learning models tend to underperform in the presence of domain shifts. Domain transfer has recently emerged as a promising approach wherein images exhibiting a domain shift are transformed into other domains for augmentation or…

图像与视频处理 · 电气工程与系统科学 2022-10-27 Weinan Song , Gaurav Fotedar , Nima Tajbakhsh , Ziheng Zhou , Lei He , Xiaowei Ding

Acoustical signal processing of directional representations of sound fields, including source, receiver, and scatterer transfer functions, are often expressed and modeled in the spherical harmonic domain (SHD). Certain such modeling…

音频与语音处理 · 电气工程与系统科学 2024-07-10 Archontis Politis

Since the number of incident energies is limited, it is difficult to directly acquire hyperspectral images (HSI) with high spatial resolution. Considering the high dimensionality and correlation of HSI, super-resolution (SR) of HSI remains…

计算机视觉与模式识别 · 计算机科学 2023-07-17 Tingting Liu , Yuan Liu , Chuncheng Zhang , Yuan Liyin , Xiubao Sui , Qian Chen

The dominant speech separation models are based on complex recurrent or convolution neural network that model speech sequences indirectly conditioning on context, such as passing information through many intermediate states in recurrent…

音频与语音处理 · 电气工程与系统科学 2020-08-17 Jingjing Chen , Qirong Mao , Dong Liu

Speaker-independent speech separation has achieved remarkable performance in recent years with the development of deep neural network (DNN). Various network architectures, from traditional convolutional neural network (CNN) and recurrent…

音频与语音处理 · 电气工程与系统科学 2022-06-17 Xue Yang , Changchun Bao

Most existing sound field reconstruction methods target point-to-region reconstruction, interpolating the Acoustic Transfer Functions (ATFs) between a fixed-position sound source and a receiver region. The applicability of these methods is…

音频与语音处理 · 电气工程与系统科学 2026-01-28 Xingyu Chen , Sipei Zhao , Fei Ma , Eva Cheng , Ian S. Burnett

This paper reformulates Transformer/Attention mechanisms in Large Language Models (LLMs) through measure theory and frequency analysis, theoretically demonstrating that hallucination is an inevitable structural limitation. The embedding…

计算与语言 · 计算机科学 2026-02-17 Kiyotaka Kasubuchi , Kazuo Fukiya

Change detection from synthetic aperture radar (SAR) imagery is a critical yet challenging task. Existing methods mainly focus on feature extraction in spatial domain, and little attention has been paid to frequency domain. Furthermore, in…

计算机视觉与模式识别 · 计算机科学 2022-02-16 Xiaofan Qu , Feng Gao , Junyu Dong , Qian Du , Heng-Chao Li

Head-related transfer functions (HRTFs) are crucial for spatial soundfield reproduction in virtual reality applications. However, obtaining personalized, high-resolution HRTFs is a time-consuming and costly task. Recently, deep…

音频与语音处理 · 电气工程与系统科学 2023-09-18 Xingyu Chen , Fei Ma , Yile Zhang , Amy Bastine , Prasanga N. Samarasinghe

Deep learning-based methods deliver state-of-the-art performance for solving inverse problems that arise in computational imaging. These methods can be broadly divided into two groups: (1) learn a network to map measurements to the signal…

图像与视频处理 · 电气工程与系统科学 2023-10-11 Nebiyou Yismaw , Ulugbek S. Kamilov , M. Salman Asif

Direct-path relative transfer function (DP-RTF) refers to the ratio between the direct-path acoustic transfer functions of two microphone channels. Though DP-RTF fully encodes the sound spatial cues and serves as a reliable localization…

声音 · 计算机科学 2022-02-17 Bing Yang , Hong Liu , Xiaofei Li
‹ 上一页 1 2 3 10 下一页 ›