中文
相关论文

相关论文: DocScanner: Robust Document Image Rectification wi…

200 篇论文

Automatic document content processing is affected by artifacts caused by the shape of the paper, non-uniform and diverse color of lighting conditions. Fully-supervised methods on real data are impossible due to the large amount of data…

计算机视觉与模式识别 · 计算机科学 2020-12-01 Sagnik Das , Hassan Ahmed Sial , Ke Ma , Ramon Baldrich , Maria Vanrell , Dimitris Samaras

The image reconstruction process in medical imaging can be treated as solving an inverse problem. The inverse problem is usually solved using time-consuming iterative algorithms with sparsity or other constraints. Recently, deep neural…

医学物理 · 物理学 2021-10-29 Jingke Zhang , Qiong He , Congzhi Wang , Hongen Liao , Jianwen Luo

Binarization of degraded document images is an elementary step in most of the problems in document image analysis domain. The paper re-visits the binarization problem by introducing an adversarial learning approach. We construct a Texture…

计算机视觉与模式识别 · 计算机科学 2019-05-02 Ankan Kumar Bhunia , Ayan Kumar Bhunia , Aneeshan Sain , Partha Pratim Roy

In this work we propose a deep learning network for deformable image registration (DIRNet). The DIRNet consists of a convolutional neural network (ConvNet) regressor, a spatial transformer, and a resampler. The ConvNet analyzes a pair of…

计算机视觉与模式识别 · 计算机科学 2017-12-08 Bob D. de Vos , Floris F. Berendsen , Max A. Viergever , Marius Staring , Ivana Išgum

A crucial problem in learning disentangled image representations is controlling the degree of disentanglement during image editing, while preserving the identity of objects. In this work, we propose a simple yet effective model with the…

机器学习 · 计算机科学 2019-12-30 Zengjie Song , Oluwasanmi Koyejo , Jiangshe Zhang

Document images are often degraded by various stains, significantly impacting their readability and hindering downstream applications such as document digitization and analysis. The absence of a comprehensive stained document dataset has…

计算机视觉与模式识别 · 计算机科学 2024-10-31 Mingxian Li , Hao Sun , Yingtie Lei , Xiaofeng Zhang , Yihang Dong , Yilin Zhou , Zimeng Li , Xuhang Chen

In this paper, a novel neural network architecture is proposed attempting to rectify text images with mild assumptions. A new dataset of text images is collected to verify our model and open to public. We explored the capability of deep…

计算机视觉与模式识别 · 计算机科学 2016-11-15 Chengzhe Yan , Jie Hu , Changshui Zhang

We introduce DocSCAN, a completely unsupervised text classification approach using Semantic Clustering by Adopting Nearest-Neighbors (SCAN). For each document, we obtain semantically informative vectors from a large pre-trained language…

计算与语言 · 计算机科学 2022-10-05 Dominik Stammbach , Elliott Ash

Text detection and recognition in natural images have long been considered as two separate tasks that are processed sequentially. Training of two tasks in a unified framework is non-trivial due to significant dif- ferences in optimisation…

计算机视觉与模式识别 · 计算机科学 2018-03-26 Tong He , Zhi Tian , Weilin Huang , Chunhua Shen , Yu Qiao , Changming Sun

In this paper we present a generalized Deep Learning-based approach for solving ill-posed large-scale inverse problems occuring in medical image reconstruction. Recently, Deep Learning methods using iterative neural networks and cascaded…

图像与视频处理 · 电气工程与系统科学 2020-08-26 Andreas Kofler , Markus Haltmeier , Tobias Schaeffter , Marc Kachelrieß , Marc Dewey , Christian Wald , Christoph Kolbitsch

Modern consumer cameras commonly employ the rolling shutter (RS) imaging mechanism, via which images are captured by scanning scenes row-by-row, resulting in RS distortion for dynamic scenes. To correct RS distortion, existing methods adopt…

计算机视觉与模式识别 · 计算机科学 2024-08-22 Wei Shang , Dongwei Ren , Wanying Zhang , Qilong Wang , Pengfei Zhu , Wangmeng Zuo

With the popularity of monocular videos generated by video sharing and live broadcasting applications, reconstructing and editing dynamic scenes in stationary monocular cameras has become a special but anticipated technology. In contrast to…

计算机视觉与模式识别 · 计算机科学 2024-02-02 Weixing Xie , Xiao Dong , Yong Yang , Qiqin Lin , Jingze Chen , Junfeng Yao , Xiaohu Guo

The advent of Multimodal Large Language Models (MLLMs) has unlocked the potential for end-to-end document parsing and translation. However, prevailing benchmarks such as OmniDocBench and DITrans are dominated by pristine scanned or…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Yongkun Du , Pinxuan Chen , Xuye Ying , Zhineng Chen

Image registration is a fundamental task in medical image analysis. Recently, deep learning based image registration methods have been extensively investigated due to their excellent performance despite the ultra-fast computational time.…

计算机视觉与模式识别 · 计算机科学 2020-08-14 Boah Kim , Dong Hwan Kim , Seong Ho Park , Jieun Kim , June-Goo Lee , Jong Chul Ye

Magnetic Resonance Imaging (MRI) is a powerful medical imaging modality, but unfortunately suffers from long scan times which, aside from increasing operational costs, can lead to image artifacts due to patient motion. Motion during the…

图像与视频处理 · 电气工程与系统科学 2023-10-02 Brett Levac , Sidharth Kumar , Ajil Jalal , Jonathan I. Tamir

In recent years, document processing has flourished and brought numerous benefits. However, there has been a significant rise in reported cases of forged document images. Specifically, recent advancements in deep neural network (DNN)…

计算机视觉与模式识别 · 计算机科学 2023-11-08 Yamato Okamoto , Osada Genki , Iu Yahiro , Rintaro Hasegawa , Peifei Zhu , Hirokatsu Kataoka

Compressed sensing (CS) is an innovative technique allowing to represent signals through a small number of their linear projections. In this paper we address the application of CS to the scenario of progressive acquisition of 2D visual…

信息论 · 计算机科学 2014-03-06 Giulio Coluccia , Enrico Magli

Most of the classical denoising methods restore clear results by selecting and averaging pixels in the noisy input. Instead of relying on hand-crafted selecting and averaging strategies, we propose to explicitly learn this process with deep…

计算机视觉与模式识别 · 计算机科学 2019-04-16 Xiangyu Xu , Muchen Li , Wenxiu Sun

Image tokenization plays a central role in modern generative modeling by mapping visual inputs into compact representations that serve as an intermediate signal between pixels and generative models. Diffusion-based decoders have recently…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Chuhan Wang , Hao Chen

This paper considers arbitrary document detection performed on a mobile device. The classical contour-based approach often fails in cases featuring occlusion, complex background, or blur. The region-based approach, which relies on the…

计算机视觉与模式识别 · 计算机科学 2021-07-02 Daniil V. Tropin , Sergey A. Ilyuhin , Dmitry P. Nikolaev , Vladimir V. Arlazarov