中文
相关论文

相关论文: NDF+: Joint Neural Directional Filtering and Diffu…

200 篇论文

Restoring degraded music signals is essential to enhance audio quality for downstream music manipulation. Recent diffusion-based music restoration methods have demonstrated impressive performance, and among them, diffusion posterior…

音频与语音处理 · 电气工程与系统科学 2023-09-14 Carlos Hernandez-Olivan , Koichi Saito , Naoki Murata , Chieh-Hsin Lai , Marco A. Martínez-Ramirez , Wei-Hsiang Liao , Yuki Mitsufuji

With the development of deep learning, speech enhancement has been greatly optimized in terms of speech quality. Previous methods typically focus on the discriminative supervised learning or generative modeling, which tends to introduce…

音频与语音处理 · 电气工程与系统科学 2025-10-31 Nan Xu , Zhaolong Huang , Xiaonan Zhi

Multi-channel speech enhancement aims to recover clean speech from noisy multi-channel recordings. Most deep learning methods employ discriminative training, which can lead to non-linear distortions from regression-based objectives,…

音频与语音处理 · 电气工程与系统科学 2026-03-26 Zhongweiyang Xu , Ashutosh Pandey , Juan Azcarreta , Zhaoheng Ni , Sanjeel Parekh , Buye Xu

Multi-frame algorithms for single-microphone speech enhancement, e.g., the multi-frame minimum variance distortionless response (MFMVDR) filter, are able to exploit speech correlation across adjacent time frames in the short-time Fourier…

音频与语音处理 · 电气工程与系统科学 2021-05-17 Marvin Tammen , Simon Doclo

The aim of this study is to implement a method to remove ambient noise in biomedical sounds captured in auscultation. We propose an incremental approach based on multichannel non-negative matrix partial co-factorization (NMPCF) for ambient…

A promising approach for multi-microphone speech separation involves two deep neural networks (DNN), where the predicted target speech from the first DNN is used to compute signal statistics for time-invariant minimum variance…

声音 · 计算机科学 2021-10-04 Zhong-Qiu Wang , Gordon Wichern , Jonathan Le Roux

We propose a spatial diffuseness feature for deep neural network (DNN)-based automatic speech recognition to improve recognition accuracy in reverberant and noisy environments. The feature is computed in real-time from multiple microphone…

计算与语言 · 计算机科学 2015-09-02 Andreas Schwarz , Christian Huemmer , Roland Maas , Walter Kellermann

Neural implicit reconstruction via volume rendering has demonstrated its effectiveness in recovering dense 3D surfaces. However, it is non-trivial to simultaneously recover meticulous geometry and preserve smoothness across regions with…

计算机视觉与模式识别 · 计算机科学 2025-04-09 Ziyu Tang , Weicai Ye , Yifan Wang , Di Huang , Hujun Bao , Tong He , Guofeng Zhang

Video denoising aims at removing noise from videos to recover clean ones. Some existing works show that optical flow can help the denoising by exploiting the additional spatial-temporal clues from nearby frames. However, the flow estimation…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Jiezhang Cao , Qin Wang , Jingyun Liang , Yulun Zhang , Kai Zhang , Radu Timofte , Luc Van Gool

FM Synthesis is a well-known algorithm used to generate complex timbre from a compact set of design primitives. Typically featuring a MIDI interface, it is usually impractical to control it from an audio source. On the other hand,…

声音 · 计算机科学 2022-08-15 Franco Caspe , Andrew McPherson , Mark Sandler

We propose a novel Neural Steering technique that adapts the target area of a spatial-aware multi-microphone sound source separation algorithm during inference without the necessity of retraining the deep neural network (DNN). To achieve…

音频与语音处理 · 电气工程与系统科学 2024-10-23 Martin Strauss , Wolfgang Mack , María Luis Valero , Okan Köpüklü

Generating diverse and high-quality 3D assets automatically poses a fundamental yet challenging task in 3D computer vision. Despite extensive efforts in 3D generation, existing optimization-based approaches struggle to produce large-scale…

计算机视觉与模式识别 · 计算机科学 2024-05-15 Ziang Cao , Fangzhou Hong , Tong Wu , Liang Pan , Ziwei Liu

Image fusion aims to generate a high-quality image from multiple images captured under varying conditions. The key problem of this task is to preserve complementary information while filtering out irrelevant information for the fused…

计算机视觉与模式识别 · 计算机科学 2023-09-04 Yuanshen Guan , Ruikang Xu , Mingde Yao , Lizhi Wang , Zhiwei Xiong

Channel estimation and beamforming play critical roles in frequency-division duplexing (FDD) massive multiple-input multiple-output (MIMO) systems. However, these two modules have been treated as two stand-alone components, which makes it…

信号处理 · 电气工程与系统科学 2021-08-04 Yifan Ma , Yifei Shen , Xianghao Yu , Jun Zhang , S. H. Song , Khaled B. Letaief

In the visual generative area, discrete diffusion models are gaining traction for their efficiency and compatibility. However, pioneered attempts still fall behind their continuous counterparts, which we attribute to noise (absorbing state)…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Tianren Ma , Xiaosong Zhang , Boyu Yang , Junlan Feng , Qixiang Ye

Finding optimal channel dimensions (i.e., the number of filters in DNN layers) is essential to design DNNs that perform well under computational resource constraints. Recent work in neural architecture search aims at automating the…

机器学习 · 计算机科学 2023-06-16 Ahmet Caner Yüzügüler , Nikolaos Dimitriadis , Pascal Frossard

Denoising diffusion models (DDMs) have recently attracted increasing attention by showing impressive synthesis quality. DDMs are built on a diffusion process that pushes data to the noise distribution and the models learn to denoise. In…

机器学习 · 计算机科学 2023-05-16 Jaemoo Choi , Yesom Park , Myungjoo Kang

Image denoising is a fundamental challenge in computer vision, with applications in photography and medical imaging. While deep learning-based methods have shown remarkable success, their reliance on specific noise distributions limits…

计算机视觉与模式识别 · 计算机科学 2025-08-28 Dongjin Kim , Jaekyun Ko , Muhammad Kashif Ali , Tae Hyun Kim

This paper proposes a deep neural network (DNN)-based multi-channel speech enhancement system in which a DNN is trained to maximize the quality of the enhanced time-domain signal. DNN-based multi-channel speech enhancement is often…

音频与语音处理 · 电气工程与系统科学 2020-02-17 Yoshiki Masuyama , Masahito Togami , Tatsuya Komatsu

This paper presents a novel machine-hearing system that exploits deep neural networks (DNNs) and head movements for robust binaural localisation of multiple sources in reverberant environments. DNNs are used to learn the relationship…

音频与语音处理 · 电气工程与系统科学 2019-04-08 Ning Ma , Tobias May , Guy J. Brown