中文
相关论文

相关论文: NDF+: Joint Neural Directional Filtering and Diffu…

200 篇论文

Neural Radiance Fields and 3D Gaussian Splatting have revolutionized 3D reconstruction and novel-view synthesis task. However, achieving photorealistic rendering from extreme novel viewpoints remains challenging, as artifacts persist across…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Jay Zhangjie Wu , Yuxuan Zhang , Haithem Turki , Xuanchi Ren , Jun Gao , Mike Zheng Shou , Sanja Fidler , Zan Gojcic , Huan Ling

Rendering dynamic reverberation in a complicated acoustic space for moving sources and listeners is challenging but crucial for enhancing user immersion in extended-reality (XR) applications. Capturing spatially varying room impulse…

音频与语音处理 · 电气工程与系统科学 2026-02-10 Orchisama Das , Gloria Dal Santo , Sebastian J. Schlecht , Vesa Valimaki , Zoran Cvetkovic

Multichannel processing is widely used for speech enhancement but several limitations appear when trying to deploy these solutions to the real-world. Distributed sensor arrays that consider several devices with a few microphones is a viable…

声音 · 计算机科学 2020-03-17 Nicolas Furnon , Romain Serizel , Irina Illina , Slim Essid

Monaural source separation is important for many real world applications. It is challenging because, with only a single channel of information available, without any constraints, an infinite number of solutions are possible. In this paper,…

声音 · 计算机科学 2015-10-02 Po-Sen Huang , Minje Kim , Mark Hasegawa-Johnson , Paris Smaragdis

Low-dose computed tomography (LDCT) reduces radiation exposure but suffers from image artifacts and loss of detail due to quantum and electronic noise, potentially impacting diagnostic accuracy. Transformer combined with diffusion models…

图像与视频处理 · 电气工程与系统科学 2025-07-01 Qiqing Liu , Guoquan Wei , Zekun Zhou , Yiyang Wen , Liu Shi , Qiegen Liu

This paper describes the practical response- and performance-aware development of online speech enhancement for an augmented reality (AR) headset that helps a user understand conversations made in real noisy echoic environments (e.g.,…

音频与语音处理 · 电气工程与系统科学 2022-07-18 Kouhei Sekiguchi , Aditya Arie Nugraha , Yicheng Du , Yoshiaki Bando , Mathieu Fontaine , Kazuyoshi Yoshii

This paper presents two single channel speech dereverberation methods to enhance the quality of speech signals that have been recorded in an enclosed space. For both methods, the room acoustics are modeled using a nonnegative approximation…

声音 · 计算机科学 2017-09-19 Nasser Mohammadiha , Simon Doclo

A neural network is essentially a high-dimensional complex mapping model by adjusting network weights for feature fitting. However, the spectral bias in network training leads to unbearable training epochs for fitting the high-frequency…

信号处理 · 电气工程与系统科学 2021-06-22 Zhi Zeng , Pengpeng Shi , Fulei Ma , Peihan Qi

This paper presents PolyDiffuse, a novel structured reconstruction algorithm that transforms visual sensor data into polygonal shapes with Diffusion Models (DM), an emerging machinery amid exploding generative AI, while formulating…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Jiacheng Chen , Ruizhi Deng , Yasutaka Furukawa

Diffusion models have demonstrated their utility as learned priors for solving various inverse problems. However, most existing approaches are limited to linear inverse problems. This paper exploits the efficient and unsupervised posterior…

图像与视频处理 · 电气工程与系统科学 2025-01-07 Mehmet Onurcan Kaya , Figen S. Oktem

Channel estimation is one of the most important parts in current mobile communication systems. Among the huge contributions in channel estimation studies, the discrete Fourier transform (DFT)-based channel estimation has attracted lots of…

信息论 · 计算机科学 2015-04-29 H. Yu , C. Yang

In this paper, we introduce a spectral-domain inverse filtering approach for single-channel speech de-reverberation using deep convolutional neural network (CNN). The main goal is to better handle realistic reverberant conditions where the…

声音 · 计算机科学 2020-10-16 Hanwook Chung , Vikrant Singh Tomar , Benoit Champagne

Deep neural networks (DNNs) have substantial computational requirements, which greatly limit their performance in resource-constrained environments. Recently, there are increasing efforts on optical neural networks and optical computing…

机器学习 · 计算机科学 2021-04-05 Yingjie Li , Ruiyang Chen , Berardi Sensale Rodriguez , Weilu Gao , Cunxi Yu

Deep neural networks (DNNs) have achieved significant success in numerous applications. The remarkable performance of DNNs is largely attributed to the availability of massive, high-quality training datasets. However, processing such…

声音 · 计算机科学 2024-07-23 Wenbo Jiang , Rui Zhang , Hongwei Li , Xiaoyuan Liu , Haomiao Yang , Shui Yu

This paper addresses the problem of speech separation and enhancement from multichannel convolutive and noisy mixtures, \emph{assuming known mixing filters}. We propose to perform the speech separation and enhancement task in the short-time…

声音 · 计算机科学 2019-01-31 Xiaofei Li , Laurent Girin , Sharon Gannot , Radu Horaud

Adapting the Diffusion Probabilistic Model (DPM) for direct image super-resolution is wasteful, given that a simple Convolutional Neural Network (CNN) can recover the main low-frequency content. Therefore, we present ResDiff, a novel…

计算机视觉与模式识别 · 计算机科学 2024-02-05 Shuyao Shang , Zhengyang Shan , Guangxing Liu , LunQian Wang , XingHua Wang , Zekai Zhang , Jinglin Zhang

This paper aims at eliminating the interfering speakers' speech, additive noise, and reverberation from the noisy multi-talker speech mixture that benefits automatic speech recognition (ASR) backend. While the recently proposed Weighted…

音频与语音处理 · 电气工程与系统科学 2020-11-19 Zhaoheng Ni , Yong Xu , Meng Yu , Bo Wu , Shixiong Zhang , Dong Yu , Michael I Mandel

In this paper, we address a statistical model extension of multichannel nonnegative matrix factorization (MNMF) for blind source separation, and we propose a new parameter update algorithm used in the sub-Gaussian model. MNMF employs…

In this paper, we propose a novel Convolutional Neural Network (CNN) structure for general-purpose multi-task learning (MTL), which enables automatic feature fusing at every layer from different tasks. This is in contrast with the most…

计算机视觉与模式识别 · 计算机科学 2019-04-08 Yuan Gao , Jiayi Ma , Mingbo Zhao , Wei Liu , Alan L. Yuille

The Reflow operation aims to straighten the inference trajectories of the rectified flow during training by constructing deterministic couplings between noises and images, thereby improving the quality of generated images in single-step or…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Jimin Dai , Jiexi Yan , Jian Yang , Lei Luo