中文
相关论文

相关论文: Seeing Straight: Document Orientation Detection fo…

200 篇论文

Motivated by the increasing popularity of transformers in computer vision, in recent times there has been a rapid development of novel architectures. While in-domain performance follows a constant, upward trend, properties like robustness…

计算机视觉与模式识别 · 计算机科学 2023-03-21 Pau de Jorge , Riccardo Volpi , Philip Torr , Gregory Rogez

Outcome Reporting Bias (ORB) poses significant threats to the validity of meta-analytic findings. It occurs when researchers selectively report outcomes based on the significance or direction of results, potentially leading to distorted…

统计方法学 · 统计学 2025-07-17 Alessandra Gaia Saracini , Leonhard Held

Compared with flatbed scanners, portable smartphones provide more convenience for physical document digitization. However, such digitized documents are often distorted due to uncontrolled physical deformations, camera positions, and…

计算机视觉与模式识别 · 计算机科学 2022-12-27 Hao Feng , Wengang Zhou , Jiajun Deng , Qi Tian , Houqiang Li

Despite powering sensitive systems like autonomous vehicles, object detection remains fairly brittle in part due to annotation errors that plague most real-world training datasets. We propose ObjectLab, a straightforward algorithm to detect…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Ulyana Tkachenko , Aditya Thyagarajan , Jonas Mueller

We apply convolutional neural networks (CNN) to the problem of image orientation detection in the context of determining the correct orientation (from 0, 90, 180, and 270 degrees) of a consumer photo. The problem is especially important for…

计算机视觉与模式识别 · 计算机科学 2023-06-02 Ujash Joshi , Michael Guerzhoy

This paper introduces a novel rotation-based framework for arbitrary-oriented text detection in natural scene images. We present the Rotation Region Proposal Networks (RRPN), which are designed to generate inclined proposals with text…

计算机视觉与模式识别 · 计算机科学 2018-10-17 Jianqi Ma , Weiyuan Shao , Hao Ye , Li Wang , Hong Wang , Yingbin Zheng , Xiangyang Xue

Practical face recognition has been studied in the past decades, but still remains an open challenge. Current prevailing approaches have already achieved substantial breakthroughs in recognition accuracy. However, their performance usually…

计算机视觉与模式识别 · 计算机科学 2015-11-03 Yandong Wen , Weiyang Liu , Meng Yang , Zhifeng Li

Recognizing text from natural images is a hot research topic in computer vision due to its various applications. Despite the enduring research of several decades on optical character recognition (OCR), recognizing texts from natural images…

计算机视觉与模式识别 · 计算机科学 2018-03-23 Zhanzhan Cheng , Yangliu Xu , Fan Bai , Yi Niu , Shiliang Pu , Shuigeng Zhou

Having a reliable accuracy score is crucial for real world applications of OCR, since such systems are judged by the number of false readings. Lexicon-based OCR systems, which deal with what is essentially a multi-class classification…

计算机视觉与模式识别 · 计算机科学 2018-07-17 Noam Mor , Lior Wolf

Document parsing is a fine-grained task where image resolution significantly impacts performance. While advanced research leveraging vision-language models benefits from high-resolution input to boost model performance, this often leads to…

Detecting oriented objects along with estimating their rotation information is one crucial step for analyzing remote sensing images. Despite that many methods proposed recently have achieved remarkable performance, most of them directly…

计算机视觉与模式识别 · 计算机科学 2021-12-02 Yanjie Wang , Xu Zou , Zhijun Zhang , Wenhui Xu , Liqun Chen , Sheng Zhong , Luxin Yan , Guodong Wang

The present work demonstrates a fast and improved technique for dewarping nonlinearly warped document images. The images are first dewarped at the page-level by estimating optimum inverse projections using curvilinear homography. The…

计算机视觉与模式识别 · 计算机科学 2020-03-17 Tanmoy Dasgupta , Nibaran Das , Mita Nasipuri

We present olmOCR 2, the latest in our family of powerful OCR systems for converting digitized print documents, like PDFs, into clean, naturally ordered plain text. olmOCR 2 is powered by olmOCR-2-7B-1025, a specialized, 7B vision language…

计算机视觉与模式识别 · 计算机科学 2025-10-23 Jake Poznanski , Luca Soldaini , Kyle Lo

We have benchmarked the maximum obtainable recognition accuracy on various word image datasets using manual segmentation and a currently available commercial OCR. We have developed a Matlab program, with graphical user interface, for…

计算机视觉与模式识别 · 计算机科学 2012-08-31 Deepak Kumar , M N Anil Prasad , A G Ramakrishnan

As the current initialization method in the state-of-the-art Stereo Visual-Inertial SLAM framework, ORB-SLAM3 has limitations. Its success depends on the performance of the pure stereo SLAM system and is based on the underlying assumption…

机器人学 · 计算机科学 2024-08-20 Han Song , Zhongche Qu , Zhi Zhang , Zihan Ye , Cong Liu

No standardized benchmark exists for evaluating OCR on food packaging, despite its critical role in automated halal food verification. Existing benchmarks target documents or scene text, missing the unique challenges of ingredient labels:…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Hasan Arief

Recent advancements in deep neural networks have markedly enhanced the performance of computer vision tasks, yet the specialized nature of these networks often necessitates extensive data and high computational power. Addressing these…

计算机视觉与模式识别 · 计算机科学 2024-01-03 Jiayou Chao , Wei Zhu

In the rapidly evolving field of optical engineering, precise alignment of multi-lens imaging systems is critical yet challenging, as even minor misalignments can significantly degrade performance. Traditional alignment methods rely on…

光学 · 物理学 2025-07-01 Tomer Slor , Dean Oren , Shira Baneth , Tom Coen , Haim Suchowski

The use of local detectors and descriptors in typical computer vision pipelines work well until variations in viewpoint and appearance change become extreme. Past research in this area has typically focused on one of two approaches to this…

计算机视觉与模式识别 · 计算机科学 2022-03-25 Udit Singh Parihar , Aniket Gujarathi , Kinal Mehta , Satyajit Tourani , Sourav Garg , Michael Milford , K. Madhava Krishna

Segmentation of a text-document into lines, words and characters, which is considered to be the crucial pre-processing stage in Optical Character Recognition (OCR) is traditionally carried out on uncompressed documents, although most of the…

计算机视觉与模式识别 · 计算机科学 2014-04-01 Mohammed Javed , P. Nagabhushan , B. B. Chaudhuri