English
Related papers

Related papers: DDI-100: Dataset for Text Detection and Recognitio…

200 papers

Text line segmentation is one of the key steps in historical document understanding. It is challenging due to the variety of fonts, contents, writing styles and the quality of documents that have degraded through the years. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Mélodie Boillet , Christopher Kermorvant , Thierry Paquet

Diffusion tensor imaging (DTI) holds significant importance in clinical diagnosis and neuroscience research. However, conventional model-based fitting methods often suffer from sensitivity to noise, leading to decreased accuracy in…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Jialong Li , Zhicheng Zhang , Yunwei Chen , Qiqi Lu , Ye Wu , Xiaoming Liu , QianJin Feng , Yanqiu Feng , Xinyuan Zhang

Historical documents encompass a wealth of cultural treasures but suffer from severe damages including character missing, paper damage, and ink erosion over time. However, existing document processing methods primarily focus on…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Zhenhua Yang , Dezhi Peng , Yongxin Shi , Yuyi Zhang , Chongyu Liu , Lianwen Jin

The possibility of carrying out a meaningful forensics analysis on printed and scanned images plays a major role in many applications. First of all, printed documents are often associated with criminal activities, such as terrorist plans,…

Computer Vision and Pattern Recognition · Computer Science 2021-02-16 Anselmo Ferreira , Ehsan Nowroozi , Mauro Barni

Recovering degraded low-resolution text images is challenging, especially for Chinese text images with complex strokes and severe degradation in real-world scenarios. Ensuring both text fidelity and style realness is crucial for…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Yuzhe Zhang , Jiawei Zhang , Hao Li , Zhouxia Wang , Luwei Hou , Dongqing Zou , Liheng Bian

Existing text recognition methods usually need large-scale training data. Most of them rely on synthetic training data due to the lack of annotated real images. However, there is a domain gap between the synthetic data and real data, which…

Computer Vision and Pattern Recognition · Computer Science 2023-03-03 Mingkun Yang , Minghui Liao , Pu Lu , Jing Wang , Shenggao Zhu , Hualin Luo , Qi Tian , Xiang Bai

Segmentation of a text-document into lines, words and characters, which is considered to be the crucial pre-processing stage in Optical Character Recognition (OCR) is traditionally carried out on uncompressed documents, although most of the…

Computer Vision and Pattern Recognition · Computer Science 2014-04-01 Mohammed Javed , P. Nagabhushan , B. B. Chaudhuri

This work evaluates six state-of-the-art deep neural network (DNN) architectures applied to the problem of enhancing camera-captured document images. The results from each network were evaluated both qualitatively and quantitatively using…

Computer Vision and Pattern Recognition · Computer Science 2021-06-30 Lucas N. Kirsten , Ricardo Piccoli , Ricardo Ribani

Imagery texts are usually organized as a hierarchy of several visual elements, i.e. characters, words, text lines and text blocks. Among these elements, character is the most basic one for various languages such as Western, Chinese,…

Computer Vision and Pattern Recognition · Computer Science 2017-08-23 Han Hu , Chengquan Zhang , Yuxuan Luo , Yuzhuo Wang , Junyu Han , Errui Ding

Document images are often degraded by various stains, significantly impacting their readability and hindering downstream applications such as document digitization and analysis. The absence of a comprehensive stained document dataset has…

Computer Vision and Pattern Recognition · Computer Science 2024-10-31 Mingxian Li , Hao Sun , Yingtie Lei , Xiaofeng Zhang , Yihang Dong , Yilin Zhou , Zimeng Li , Xuhang Chen

Multilingual document understanding remains limited for low-resource languages due to scarce training data and model-based annotation pipelines that perpetuate existing biases. We introduce DocAtlas, a framework that constructs…

Document Visual Question Answering (VQA) aims to understand visually-rich documents to answer questions in natural language, which is an emerging research topic for both Natural Language Processing and Computer Vision. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2023-05-05 Fengbin Zhu , Wenqiang Lei , Fuli Feng , Chao Wang , Haozhou Zhang , Tat-Seng Chua

Despite significant progress on current state-of-the-art image generation models, synthesis of document images containing multiple and complex object layouts is a challenging task. This paper presents a novel approach, called DocSynth, to…

Computer Vision and Pattern Recognition · Computer Science 2021-07-07 Sanket Biswas , Pau Riba , Josep Lladós , Umapada Pal

Capturing images is a key part of automation for high-level tasks such as scene text recognition. Low-light conditions pose a challenge for high-level perception stacks, which are often optimized on well-lit, artifact-free images.…

Image and Video Processing · Electrical Eng. & Systems 2023-11-01 Cindy M. Nguyen , Eric R. Chan , Alexander W. Bergman , Gordon Wetzstein

Shadows often occur when we capture the documents with casual equipment, which influences the visual quality and readability of the digital copies. Different from the algorithms for natural shadow removal, the algorithms in document shadow…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Zinuo Li , Xuhang Chen , Chi-Man Pun , Xiaodong Cun

Multi-focus image fusion, a technique to generate an all-in-focus image from two or more partially-focused source images, can benefit many computer vision tasks. However, currently there is no large and realistic dataset to perform…

Computer Vision and Pattern Recognition · Computer Science 2020-08-31 Juncheng Zhang , Qingmin Liao , Shaojun Liu , Haoyu Ma , Wenming Yang , Jing-Hao Xue

The examination of the musculoskeletal system in dogs is a challenging task in veterinary practice. In this work, a novel method has been developed that enables efficient documentation of a dog's condition through a visual representation.…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Martin Thißen , Thi Ngoc Diep Tran , Ben Joel Schönbein , Ute Trapp , Barbara Esteve Ratsch , Beate Egner , Romana Piat , Elke Hergenröther

Document images are now widely captured by handheld devices such as mobile phones. The OCR performance on these images are largely affected due to geometric distortion of the document paper, diverse camera positions and complex backgrounds.…

Computer Vision and Pattern Recognition · Computer Science 2022-03-22 Guo-Wang Xie , Fei Yin , Xu-Yao Zhang , Cheng-Lin Liu

Automating the annotation of scanned documents is challenging, requiring a balance between computational efficiency and accuracy. DocParseNet addresses this by combining deep learning and multi-modal learning to process both text and visual…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Ahmad Mohammadshirazi , Ali Nosrati Firoozsalari , Mengxi Zhou , Dheeraj Kulshrestha , Rajiv Ramnath

Real-world applications of computer vision in the humanities require algorithms to be robust against artistic abstraction, peripheral objects, and subtle differences between fine-grained target classes. Existing datasets provide…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Mathias Zinnen , Prathmesh Madhu , Inger Leemans , Peter Bell , Azhar Hussian , Hang Tran , Ali Hürriyetoğlu , Andreas Maier , Vincent Christlein