中文
相关论文

相关论文: UVDoc: Neural Grid-based Document Unwarping

200 篇论文

Document dewarping is crucial for many applications. However, existing learning-based methods rely heavily on supervised regression with annotated data without fully leveraging the inherent geometric properties of physical documents. Our…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Chaoyun Wang , I-Chao Shen , Takeo Igarashi , Caigui Jiang

The task of predicting smooth and edge-consistent depth maps is notoriously difficult for single image depth estimation. This paper proposes a novel Bilateral Grid based 3D convolutional neural network, dubbed as 3DBG-UNet, that…

计算机视觉与模式识别 · 计算机科学 2021-05-24 Mansi Sharma , Abheesht Sharma , Kadvekar Rohit Tushar , Avinash Panneer

We introduce Neural Deformation Graphs for globally-consistent deformation tracking and 3D reconstruction of non-rigid objects. Specifically, we implicitly model a deformation graph via a deep neural network. This neural deformation graph…

计算机视觉与模式识别 · 计算机科学 2020-12-04 Aljaž Božič , Pablo Palafox , Michael Zollhöfer , Justus Thies , Angela Dai , Matthias Nießner

We introduce neural dual contouring (NDC), a new data-driven approach to mesh reconstruction based on dual contouring (DC). Like traditional DC, it produces exactly one vertex per grid cell and one quad for each grid edge intersection, a…

计算机视觉与模式识别 · 计算机科学 2022-06-02 Zhiqin Chen , Andrea Tagliasacchi , Thomas Funkhouser , Hao Zhang

Recently there has been a significant effort to automate UV mapping, the process of mapping 3D-dimensional surfaces to the UV space while minimizing distortion and seam length. Although state-of-the-art methods, Autocuts and OptCuts,…

图形学 · 计算机科学 2020-12-04 Fatemeh Teimury , Bruno Roy , Juan Sebastián Casallas , David MacDonald , Mark Coates

We propose Universal Document Processing (UDOP), a foundation Document AI model which unifies text, image, and layout modalities together with varied task formats, including document understanding and generation. UDOP leverages the spatial…

计算机视觉与模式识别 · 计算机科学 2023-03-14 Zineng Tang , Ziyi Yang , Guoxin Wang , Yuwei Fang , Yang Liu , Chenguang Zhu , Michael Zeng , Cha Zhang , Mohit Bansal

In modern computer vision, images are typically represented as a fixed uniform grid with some stride and processed via a deep convolutional neural network. We argue that deforming the grid to better align with the high-frequency image…

计算机视觉与模式识别 · 计算机科学 2023-04-05 Jun Gao , Zian Wang , Jinchen Xuan , Sanja Fidler

We introduce a method for fast estimation of data-adapted, spatio-temporally dependent regularization parameter-maps for variational image reconstruction, focusing on total variation (TV)-minimization. Our approach is inspired by recent…

Learning to reconstruct depths in a single image by watching unlabeled videos via deep convolutional network (DCN) is attracting significant attention in recent years. In this paper, we introduce a surface normal representation for…

计算机视觉与模式识别 · 计算机科学 2017-11-13 Zhenheng Yang , Peng Wang , Wei Xu , Liang Zhao , Ramakant Nevatia

Deformable image registration is a fundamental task in medical image analysis, aiming to establish a dense and non-linear correspondence between a pair of images. Previous deep-learning studies usually employ supervised neural networks to…

计算机视觉与模式识别 · 计算机科学 2018-09-11 Jun Zhang

Although deep convolutional neural network has been proved to efficiently eliminate coding artifacts caused by the coarse quantization of traditional codec, it's difficult to train any neural network in front of the encoder for gradient's…

计算机视觉与模式识别 · 计算机科学 2018-01-17 Lijun Zhao , Huihui Bai , Anhong Wang , Yao Zhao

Map-to-map matching is a critical task for aligning spatial data across heterogeneous sources, yet it remains challenging due to the lack of ground truth correspondences, sparse node features, and scalability demands. In this paper, we…

机器学习 · 计算机科学 2026-01-21 Chaolong Ying , Yinan Zhang , Lei Zhang , Jiazhuang Wang , Shujun Jia , Tianshu Yu

Visual document understanding (VDU) has rapidly advanced with the development of powerful multi-modal language models. However, these models typically require extensive document pre-training data to learn intermediate representations and…

计算机视觉与模式识别 · 计算机科学 2024-11-06 Souhail Bakkali , Sanket Biswas , Zuheng Ming , Mickaël Coustaty , Marçal Rusiñol , Oriol Ramos Terrades , Josep Lladós

The present work demonstrates a fast and improved technique for dewarping nonlinearly warped document images. The images are first dewarped at the page-level by estimating optimum inverse projections using curvilinear homography. The…

计算机视觉与模式识别 · 计算机科学 2020-03-17 Tanmoy Dasgupta , Nibaran Das , Mita Nasipuri

Surface parameterization is a fundamental geometry processing problem with rich downstream applications. Traditional approaches are designed to operate on well-behaved mesh models with high-quality triangulations that are laboriously…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Qijian Zhang , Junhui Hou , Ying He

Presentation of folded documents is not an uncommon case in modern society. Digitizing such documents by capturing them with a smartphone camera can be tricky since a crease can divide the document contents into separate planes. To unfold…

计算机视觉与模式识别 · 计算机科学 2024-08-13 A. M. Ershov , D. V. Tropin , E. E. Limonova , D. P. Nikolaev , V. V. Arlazarov

Recent grid-based document representations like BERTgrid allow the simultaneous encoding of the textual and layout information of a document in a 2D feature map so that state-of-the-art image segmentation and/or object detection models can…

计算与语言 · 计算机科学 2021-05-26 Weihong Lin , Qifang Gao , Lei Sun , Zhuoyao Zhong , Kai Hu , Qin Ren , Qiang Huo

Background subtraction is a fundamental task in computer vision with numerous real-world applications, ranging from object tracking to video surveillance. Dynamic backgrounds poses a significant challenge here. Supervised deep…

计算机视觉与模式识别 · 计算机科学 2023-03-07 Fateme Bahri , Nilanjan Ray

Image quality degradation caused by raindrops is one of the most important but challenging problems that reduce the performance of vision systems. Most existing raindrop removal algorithms are based on a supervised learning method using…

计算机视觉与模式识别 · 计算机科学 2024-12-05 Huijiao Wang , Shenghao Zhao , Lei Yu , Xulei Yang

In this paper, we present NeuralReshaper, a novel method for semantic reshaping of human bodies in single images using deep generative networks. To achieve globally coherent reshaping effects, our approach follows a fit-then-reshape…

计算机视觉与模式识别 · 计算机科学 2023-08-15 Beijia Chen , Yuefan Shen , Hongbo Fu , Xiang Chen , Kun Zhou , Youyi Zheng