中文
相关论文

相关论文: Digitizing Historical Balance Sheet Data: A Practi…

200 篇论文

Commercial OCR packages work best with high-quality scanned images. They often produce poor results when the image is degraded, either because the original itself was poor quality, or because of excessive photocopying. The ability to…

数字图书馆 · 计算机科学 2007-05-23 Roger T. Hartley , Kathleen Crumpton

The problem of optical character recognition, OCR, has been widely discussed in the literature. Having a hand-written text, the program aims at recognizing the text. Even though there are several approaches to this issue, it is still an…

计算机视觉与模式识别 · 计算机科学 2014-11-07 Wei Wang

Spurious credit card transactions are a significant source of financial losses and urge the development of accurate fraud detection algorithms. In this paper, we use machine learning strategies for such an aim. First, we apply a mixed…

机器学习 · 计算机科学 2021-12-07 Daniel H. M. de Souza , Claudio J. Bordin

We explore how multimodal Large Language Models (mLLMs) can help researchers transcribe historical documents, extract relevant historical information, and construct datasets from historical sources. Specifically, we investigate the…

计算与语言 · 计算机科学 2025-04-02 Gavin Greif , Niclas Griesshaber , Robin Greif

With the rapid development of OCR technology, mixed-scene text recognition has become a key technical challenge. Although deep learning models have achieved significant results in specific scenarios, their generality and stability still…

计算机视觉与模式识别 · 计算机科学 2025-05-12 Da Chang , Yu Li

Line Chart Data Extraction is a natural extension of Optical Character Recognition where the objective is to recover the underlying numerical information a chart image represents. Some recent works such as ChartOCR approach this problem…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Shufan Li , Congxi Lu , Linkai Li , Haoshuai Zhou

Optical Character Recognition (OCR) technology has revolutionized the digitization of printed text, enabling efficient data extraction and analysis across various domains. Just like Machine Translation systems, OCR systems are prone to…

计算与语言 · 计算机科学 2025-02-25 Harshvivek Kashid , Pushpak Bhattacharyya

Post-OCR processing has significantly improved over the past few years. However, these have been primarily beneficial for texts consisting of natural, alphabetical words, as opposed to documents of numerical nature such as invoices,…

计算与语言 · 计算机科学 2023-07-04 Arthur Hemmer , Jérôme Brachat , Mickaël Coustaty , Jean-Marc Ogier

Paper-intensive industries like insurance, law, and government have long leveraged optical character recognition (OCR) to automatically transcribe hordes of scanned documents into text strings for downstream processing. Even in 2019, there…

计算机视觉与模式识别 · 计算机科学 2020-01-17 W. Ronny Huang , Yike Qi , Qianqian Li , Jonathan Degange

Extracting the relevant information out of a large number of documents is a challenging and tedious task. The quality of results generated by the traditionally available full-text search engine and text-based image retrieval systems is not…

信息检索 · 计算机科学 2022-12-05 Riya Gupta , C. V. Jawahar

This research paper introduces a novel word-level Optical Character Recognition (OCR) model specifically designed for digital Urdu text, leveraging transformer-based architectures and attention mechanisms to address the distinct challenges…

计算机视觉与模式识别 · 计算机科学 2024-09-02 Ahmed Mustafa , Muhammad Tahir Rafique , Muhammad Ijlal Baig , Hasan Sajid , Muhammad Jawad Khan , Karam Dad Kallu

We present a framework to generate synthetic historical documents with precise ground truth using nothing more than a collection of unlabeled historical images. Obtaining large labeled datasets is often the limiting factor to effectively…

计算机视觉与模式识别 · 计算机科学 2021-05-18 Lars Vögtlin , Manuel Drazyk , Vinaychandran Pondenkandath , Michele Alberti , Rolf Ingold

Document parsing is a core task in document intelligence, supporting applications such as information extraction, retrieval-augmented generation, and automated document analysis. However, real-world documents often feature complex layouts…

Credit scoring is vital in the financial industry, assessing the risk of lending to credit card applicants. Traditional credit scoring methods face challenges with large datasets and data imbalance between creditworthy and non-creditworthy…

计算工程、金融与科学 · 计算机科学 2024-09-26 Kejian Tong , Zonglin Han , Yanxin Shen , Yujian Long , Yijing Wei

Increased availability of electronic health records (EHR) has enabled researchers to study various medical questions. Cohort selection for the hypothesis under investigation is one of the main consideration for EHR analysis. For uncommon…

机器学习 · 计算机科学 2020-05-14 Mohamed Ghalwash , Zijun Yao , Prithwish Chakrabotry , James Codella , Daby Sow

Standard OCR is a well-researched topic of computer vision and can be considered solved for machine-printed text. However, when applied to unconstrained images, the recognition rates drop drastically. Therefore, the employment of object…

计算机视觉与模式识别 · 计算机科学 2013-04-29 Albert Kavelar , Sebastian Zambanini , Martin Kampel

In the proposed study, we describe the possibility of automated dataset collection using an articulated robot. The proposed technology reduces the number of pixel errors on a polygonal dataset and the time spent on manual labeling of 2D…

机器人学 · 计算机科学 2021-08-06 Valery Ilin , Ivan Kalinov , Pavel Karpyshev , Dzmitry Tsetserukou

Substantial amounts of work are required to clean large collections of digitized books for NLP analysis, both because of the presence of errors in the scanned text and the presence of duplicate volumes in the corpora. In this paper, we…

计算与语言 · 计算机科学 2021-10-25 Allen Kim , Charuta Pethe , Naoya Inoue , Steve Skiena

Optical Character Recognition (OCR) is an established task with the objective of identifying the text present in an image. While many off-the-shelf OCR models exist, they are often trained for either scientific (e.g., formulae) or generic…

计算与语言 · 计算机科学 2024-03-26 Nan Zhang , Connor Heaton , Sean Timothy Okonsky , Prasenjit Mitra , Hilal Ezgi Toraman

Many languages have vast amounts of handwritten texts, such as ancient scripts about folktale stories and historical narratives or contemporary documents and letters. Digitization of those texts has various applications, such as daily…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Ameer Majeed , Hossein Hassani