中文
相关论文

相关论文: McNet: Fuse Multiple Cues for Multichannel Speech …

200 篇论文

A typical neural speech enhancement (SE) approach mainly handles speech and noise mixtures, which is not optimal for singing voice enhancement scenarios. Music source separation (MSS) models treat vocals and various accompaniment components…

声音 · 计算机科学 2023-10-09 Weiming Xu , Zhouxuan Chen , Zhili Tan , Shubo Lv , Runduo Han , Wenjiang Zhou , Weifeng Zhao , Lei Xie

The capability of the human to pay attention to both coarse and fine-grained regions has been applied to computer vision tasks. Motivated by that, we propose a collaborative learning framework in the complex domain for monaural noise…

声音 · 计算机科学 2021-06-23 Andong Li , Chengshi Zheng , Lu Zhang , Xiaodong Li

Spectrograms have been widely used in Convolutional Neural Networks based schemes for acoustic scene classification, such as the STFT spectrogram and the MFCC spectrogram, etc. They have different time-frequency characteristics,…

计算机视觉与模式识别 · 计算机科学 2018-09-06 Weiping Zheng , Zhenyao Mo , Xiaotao Xing , Gansen Zhao

Single-channel speech enhancement algorithms are often used in resource-constrained embedded devices, where low latency and low complexity designs gain more importance. In recent years, researchers have proposed a wide variety of novel…

音频与语音处理 · 电气工程与系统科学 2026-04-29 Nicolás Arrieta Larraza , Niels de Koeijer

Recent single-channel speech enhancement methods usually convert waveform to the time-frequency domain and use magnitude/complex spectrum as the optimizing target. However, both magnitude-spectrum-based methods and complex-spectrum-based…

声音 · 计算机科学 2021-10-13 Wenxin Tai , Jiajia Li , Yixiang Wang , Tian Lan , Qiao Liu

For the difficulty and large computational complexity of modeling more frequency bands, full-band speech enhancement based on deep neural networks is still challenging. Previous studies usually adopt compressed full-band speech features in…

声音 · 计算机科学 2022-08-02 Guochen Yu , Yuansheng Guan , Weixin Meng , Chengshi Zheng , Hui Wang

Diffusion model, as a new generative model which is very popular in image generation and audio synthesis, is rarely used in speech enhancement. In this paper, we use the diffusion model as a module for stochastic refinement. We propose…

声音 · 计算机科学 2022-11-01 Zhibin Qiu , Mengfan Fu , Yinfeng Yu , LiLi Yin , Fuchun Sun , Hao Huang

Multi-modal fusion is proven to be an effective method to improve the accuracy and robustness of speaker tracking, especially in complex scenarios. However, how to combine the heterogeneous information and exploit the complementarity of…

计算机视觉与模式识别 · 计算机科学 2021-12-15 Yidi Li , Hong Liu , Hao Tang

This paper proposes the MBURST, a novel multimodal solution for audio-visual speech enhancements that consider the most recent neurological discoveries regarding pyramidal cells of the prefrontal cortex and other brain regions. The…

声音 · 计算机科学 2024-09-10 Mohsin Raza , Leandro A. Passos , Ahmed Khubaib , Ahsan Adeel

As wireless communication systems evolve, automatic modulation recognition (AMR) plays a key role in improving spectrum efficiency, especially in cognitive radio systems. Traditional AMR methods face challenges in complex, noisy…

信号处理 · 电气工程与系统科学 2025-10-22 Wangye Jiang , Haoming Yang , Xinyu Lu , Mingyuan Wang , Huimei Sun , Jingya Zhang

Recently, speech enhancement technologies that are based on deep learning have received considerable research attention. If the spatial information in microphone signals is exploited, microphone arrays can be advantageous under some adverse…

音频与语音处理 · 电气工程与系统科学 2022-07-19 Yicheng Hsu , Yonghan Lee , Mingsian R. Bai

Deep learning models, such as the fully convolutional network (FCN), have been widely used in 3D biomedical segmentation and achieved state-of-the-art performance. Multiple modalities are often used for disease diagnosis and quantification.…

图像与视频处理 · 电气工程与系统科学 2019-08-23 Yu Chen , Jiawei Chen , Dong Wei , Yuexiang Li , Yefeng Zheng

Given a video and a linguistic query, video moment retrieval and highlight detection (MR&HD) aim to locate all the relevant spans while simultaneously predicting saliency scores. Most existing methods utilize RGB images as input,…

计算机视觉与模式识别 · 计算机科学 2025-01-22 Yifang Xu , Yunzhuo Sun , Benxiang Zhai , Zien Xie , Youyao Jia , Sidan Du

Most of the current deep learning-based approaches for speech enhancement only operate in the spectrogram or waveform domain. Although a cross-domain transformer combining waveform- and spectrogram-domain inputs has been proposed, its…

声音 · 计算机科学 2023-10-31 Jialu Li , Junhui Li , Pu Wang , Youshan Zhang

Image classification models often demonstrate unstable performance in real-world applications due to variations in image information, driven by differing visual perspectives of subject objects and lighting discrepancies. To mitigate these…

计算机视觉与模式识别 · 计算机科学 2024-07-29 Yuze Zheng , Zixuan Li , Xiangxian Li , Jinxing Liu , Yuqing Wang , Xiangxu Meng , Lei Meng

The extraction of a desired speech signal from a noisy environment has become a challenging issue. In the recent years, the scientific community has particularly focused on multichannel techniques which are dealt with in this review. In…

声音 · 计算机科学 2013-01-01 Adel Hidri , Souad Meddeb , Hamid Amiri

Multi-channel acoustic signal processing is a well-established and powerful tool to exploit the spatial diversity between a target signal and non-target or noise sources for signal enhancement. However, the textbook solutions for optimal…

音频与语音处理 · 电气工程与系统科学 2025-01-14 Reinhold Haeb-Umbach , Tomohiro Nakatani , Marc Delcroix , Christoph Boeddeker , Tsubasa Ochiai

While traditional statistical signal processing model-based methods can derive the optimal estimators relying on specific statistical assumptions, current learning-based methods further promote the performance upper bound via deep neural…

声音 · 计算机科学 2022-03-17 Andong Li , Chengshi Zheng , Ziyang Zhang , Xiaodong Li

This paper investigates the optimal selection and fusion of feature encoders across multiple modalities and combines these in one neural network to improve sentiment detection. We compare different fusion methods and examine the impact of…

计算与语言 · 计算机科学 2024-06-04 Zehui Wu , Ziwei Gong , Jaywon Koo , Julia Hirschberg

The low-level spatial detail information and high-level semantic abstract information are both essential to the semantic segmentation task. The features extracted by the deep network can obtain rich semantic information, while a lot of…

计算机视觉与模式识别 · 计算机科学 2023-12-13 Xiaojie Fang , Xingguo Song , Xiangyin Meng , Xu Fang , Sheng Jin