中文
相关论文

相关论文: DocDeshadower: Frequency-Aware Transformer for Doc…

200 篇论文

Although diffusion models are rising as a powerful solution for blind face restoration, they are criticized for two problems: 1) slow training and inference speed, and 2) failure in preserving identity and recovering fine-grained facial…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Yunqi Miao , Jiankang Deng , Jungong Han

Recovering noise-covered details from low-light images is challenging, and the results given by previous methods leave room for improvement. Recent diffusion models show realistic and detailed image generation through a sequence of…

计算机视觉与模式识别 · 计算机科学 2023-05-18 Dewei Zhou , Zongxin Yang , Yi Yang

A significant volume of analog information, i.e., documents and images, have been digitized in the form of scanned copies for storing, sharing, and/or analyzing in the digital world. However, the quality of such contents is severely…

计算机视觉与模式识别 · 计算机科学 2024-02-09 Junghun Cha , Ali Haider , Seoyun Yang , Hoeyeong Jin , Subin Yang , A. F. M. Shahab Uddin , Jaehyoung Kim , Soo Ye Kim , Sung-Ho Bae

Mesh denoising, aimed at removing noise from input meshes while preserving their feature structures, is a practical yet challenging task. Despite the remarkable progress in learning-based mesh denoising methodologies in recent years, their…

计算机视觉与模式识别 · 计算机科学 2024-05-13 Wenbo Zhao , Xianming Liu , Deming Zhai , Junjun Jiang , Xiangyang Ji

For capturing colored document images, e.g. posters and magazines, it is common that multiple degradations such as shadows, wrinkles, etc., are simultaneously introduced due to external factors. Restoring multi-degraded colored document…

计算机视觉与模式识别 · 计算机科学 2023-10-30 Chaowei Liu , Jichun Li , Yihua Teng , Chaoqun Wang , Nuo Xu , Jihao Wu , Dandan Tu

The depth completion task is a critical problem in autonomous driving, involving the generation of dense depth maps from sparse depth maps and RGB images. Most existing methods employ a spatial propagation network to iteratively refine the…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Ming Yuan , Chuang Zhang , Lei He , Qing Xu , Jianqiang Wang

Digital versions of real-world text documents often suffer from issues like environmental corrosion of the original document, low-quality scanning, or human interference. Existing document restoration and inpainting methods typically…

计算机视觉与模式识别 · 计算机科学 2025-03-07 Wanglong Lu , Lingming Su , Jingjing Zheng , Vinícius Veloso de Melo , Farzaneh Shoeleh , John Hawkin , Terrence Tricco , Hanli Zhao , Xianta Jiang

Document images are often degraded by various stains, significantly impacting their readability and hindering downstream applications such as document digitization and analysis. The absence of a comprehensive stained document dataset has…

计算机视觉与模式识别 · 计算机科学 2024-10-31 Mingxian Li , Hao Sun , Yingtie Lei , Xiaofeng Zhang , Yihang Dong , Yilin Zhou , Zimeng Li , Xuhang Chen

Diffusion transformers have demonstrated remarkable generation quality, albeit requiring longer training iterations and numerous inference steps. In each denoising step, diffusion transformers encode the noisy inputs to extract the…

计算机视觉与模式识别 · 计算机科学 2025-04-10 Shuai Wang , Zhi Tian , Weilin Huang , Limin Wang

Shadow removal aims at restoring the image content within shadow regions, pursuing a uniform distribution of illumination that is consistent between shadow and non-shadow regions. {Comparing to other image restoration tasks, there are two…

计算机视觉与模式识别 · 计算机科学 2024-10-07 Laniqng Guo , Chong Wang , Yufei Wang , Yi Yu , Siyu Huang , Wenhan Yang , Alex C. Kot , Bihan Wen

In this work, we propose a novel deformable convolutional pyramid network for unsupervised image registration. Specifically, the proposed network enhances the traditional pyramid network by adding an additional shared auxiliary decoder for…

计算机视觉与模式识别 · 计算机科学 2024-08-13 Hongchao Zhou , Shunbo Hu

Recent advancements in deep learning have yielded promising results for the image shadow removal task. However, most existing methods rely on binary pre-generated shadow masks. The binary nature of such masks could potentially lead to…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Xinrui Wang , Lanqing Guo , Xiyu Wang , Siyu Huang , Bihan Wen

There is a recent trend in the LiDAR perception field towards unifying multiple tasks in a single strong network with improved performance, as opposed to using separate networks for each task. In this paper, we introduce a new LiDAR…

计算机视觉与模式识别 · 计算机科学 2024-03-05 Zixiang Zhou , Dongqiangzi Ye , Weijia Chen , Yufei Xie , Yu Wang , Panqu Wang , Hassan Foroosh

Detecting and grounding multi-modal media manipulation (DGM^4) has become increasingly crucial due to the widespread dissemination of face forgery and text misinformation. In this paper, we present the Unified Frequency-Assisted transFormer…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Huan Liu , Zichang Tan , Qiang Chen , Yunchao Wei , Yao Zhao , Jingdong Wang

Image colorization is a challenging problem due to multi-modal uncertainty and high ill-posedness. Directly training a deep neural network usually leads to incorrect semantic colors and low color richness. While transformer-based methods…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Xiaoyang Kang , Tao Yang , Wenqi Ouyang , Peiran Ren , Lingzhi Li , Xuansong Xie

In winter scenes, the degradation of images taken under snow can be pretty complex, where the spatial distribution of snowy degradation is varied from image to image. Recent methods adopt deep neural networks to directly recover clean…

计算机视觉与模式识别 · 计算机科学 2022-07-13 Tian Ye , Sixiang Chen , Yun Liu , Yi Ye , Erkang Chen

Hyperspectral image classification (HSIC) has gained significant attention because of its potential in analyzing high-dimensional data with rich spectral and spatial information. In this work, we propose the Differential Spatial-Spectral…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Muhammad Ahmad , Manuel Mazzara , Salvatore Distefano , Adil Mehmood Khan , Silvia Liberata Ullo

This paper introduces a new Transformer, called MS$^2$Dformer, that can be used as a generalized backbone for multi-modal sequence spammer detection. Spammer detection is a complex multi-modal task, thus the challenges of applying…

机器学习 · 计算机科学 2025-02-25 Zhou Yang , Yucai Pang , Hongbo Yin , Yunpeng Xiao

Image editing techniques have rapidly advanced, facilitating both innovative use cases and malicious manipulation of digital images. Deep learning-based methods have recently achieved high accuracy in pixel-level forgery localization, yet…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Ju-Hyeon Nam , Dong-Hyun Moon , Sang-Chul Lee

Capturing images is a key part of automation for high-level tasks such as scene text recognition. Low-light conditions pose a challenge for high-level perception stacks, which are often optimized on well-lit, artifact-free images.…

图像与视频处理 · 电气工程与系统科学 2023-11-01 Cindy M. Nguyen , Eric R. Chan , Alexander W. Bergman , Gordon Wetzstein