English

Learning Spatial and Spatio-Temporal Pixel Aggregations for Image and Video Denoising

Computer Vision and Pattern Recognition 2021-02-03 v1 Artificial Intelligence

Abstract

Existing denoising methods typically restore clear results by aggregating pixels from the noisy input. Instead of relying on hand-crafted aggregation schemes, we propose to explicitly learn this process with deep neural networks. We present a spatial pixel aggregation network and learn the pixel sampling and averaging strategies for image denoising. The proposed model naturally adapts to image structures and can effectively improve the denoised results. Furthermore, we develop a spatio-temporal pixel aggregation network for video denoising to efficiently sample pixels across the spatio-temporal space. Our method is able to solve the misalignment issues caused by large motion in dynamic scenes. In addition, we introduce a new regularization term for effectively training the proposed video denoising model. We present extensive analysis of the proposed method and demonstrate that our model performs favorably against the state-of-the-art image and video denoising approaches on both synthetic and real-world data.

Keywords

Cite

@article{arxiv.2101.10760,
  title  = {Learning Spatial and Spatio-Temporal Pixel Aggregations for Image and Video Denoising},
  author = {Xiangyu Xu and Muchen Li and Wenxiu Sun and Ming-Hsuan Yang},
  journal= {arXiv preprint arXiv:2101.10760},
  year   = {2021}
}

Comments

Project page: https://sites.google.com/view/xiangyuxu/denoise_stpan. arXiv admin note: substantial text overlap with arXiv:1904.06903

R2 v1 2026-06-23T22:32:35.830Z