中文
相关论文

相关论文: Ancient but Digitized: Developing Handwritten Opti…

200 篇论文

We propose a post-OCR text correction approach for digitising texts in Romanised Sanskrit. Owing to the lack of resources our approach uses OCR models trained for other languages written in Roman. Currently, there exists no dataset…

计算与语言 · 计算机科学 2018-09-10 Amrith Krishna , Bodhisattwa Prasad Majumder , Rajesh Shreedhar Bhat , Pawan Goyal

Digit, letter and word recognition for a particular script has various applications in todays commercial contexts. Nevertheless, only a limited number of relevant studies have dealt with Persian scripts. In this paper, deep neural networks…

计算机视觉与模式识别 · 计算机科学 2020-11-17 Mehdi Bonyani , Simindokht Jahangard , Morteza Daneshmand

Given the ubiquity of handwritten documents in human transactions, Optical Character Recognition (OCR) of documents have invaluable practical worth. Optical character recognition is a science that enables to translate various types of…

计算机视觉与模式识别 · 计算机科学 2020-01-03 Jamshed Memon , Maira Sami , Rizwan Ahmed Khan

The ultimate aim of handwriting recognition is to make computers able to read and/or authenticate human written texts, with a performance comparable to or even better than that of humans. Reading means that the computer is given a piece of…

计算机视觉与模式识别 · 计算机科学 2012-06-26 Manal A. Abdullah , Lulwah M. Al-Harigy , Hanadi H. Al-Fraidi

Optical character recognition (OCR) methods have been applied to diverse tasks, e.g., street view text recognition and document analysis. Recently, zero-shot OCR has piqued the interest of the research community because it considers a…

计算机视觉与模式识别 · 计算机科学 2023-08-02 Xiaolei Diao , Daqian Shi , Jian Li , Lida Shi , Mingzhe Yue , Ruihua Qi , Chuntao Li , Hao Xu

There has been recent interest in improving optical character recognition (OCR) for endangered languages, particularly because a large number of documents and books in these languages are not in machine-readable formats. The performance of…

计算与语言 · 计算机科学 2023-02-28 Shruti Rijhwani , Daisy Rosenblum , Michayla King , Antonios Anastasopoulos , Graham Neubig

The inherent complexities of Arabic script; its cursive nature, diacritical marks (tashkeel), and varied typography, pose persistent challenges for Optical Character Recognition (OCR). We present Qari-OCR, a series of vision-language models…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Ahmed Wasfy , Omer Nacar , Abdelakreem Elkhateb , Mahmoud Reda , Omar Elshehy , Adel Ammar , Wadii Boulila

Character recognition is the fundamental part of an optical character recognition (OCR) system. Word recognition, sentence transcription, document digitization, and language processing are some of the higher-order activities that can be…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Mirza Raquib , Asif Pervez Polok , Kedar Nath Biswas , Farida Siddiqi Prity , Saydul Akbar Murad , Nick Rahimi

Our study utilizes deep learning methods for the automated transcription of late nineteenth- and early twentieth-century periodicals written in Arabic script Ottoman Turkish (OT) using the Transkribus platform. We discuss the historical…

计算与语言 · 计算机科学 2020-11-03 Suphan Kirmizialtin , David Wrisley

Extracting Handwritten text is one of the most important components of digitizing information and making it available for large scale setting. Handwriting Optical Character Reader (OCR) is a research problem in computer vision and natural…

计算机视觉与模式识别 · 计算机科学 2022-06-10 Mohammad Daniyal Shaiq , Musa Dildar Ahmed Cheema , Ali Kamal

HTR models development has become a conventional step for digital humanities projects. The performance of these models, often quite high, relies on manual transcription and numerous handwritten documents. Although the method has proven…

计算机视觉与模式识别 · 计算机科学 2022-11-30 Lucas Noëmie , Clément Salah , Chahan Vidal-Gorène

The project comes with the technique of OCR (Optical Character Recognition) which includes various research sides of computer science. The project is to take a picture of a character and process it up to recognize the image of that…

计算机视觉与模式识别 · 计算机科学 2021-11-11 Arkaprabha Basu , M. Sathya

Objective of the current work is to develop an Optical Character Recognition (OCR) engine for information Just In Time (iJIT) system that can be used for recognition of handwritten textual annotations of lower case Roman script. Tesseract…

计算机视觉与模式识别 · 计算机科学 2010-03-31 Sandip Rakshit , Subhadip Basu , Hisashi Ikeda

Handwriting analysis is still an important application in machine learning. A basic requirement for any handwriting recognition application is the availability of comprehensive datasets. Standard labelled datasets play a significant role in…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Pourya Jafarzadeh , Padideh Choobdar , Vahid Mohammadi Safarzadeh

This research paper introduces a novel word-level Optical Character Recognition (OCR) model specifically designed for digital Urdu text, leveraging transformer-based architectures and attention mechanisms to address the distinct challenges…

计算机视觉与模式识别 · 计算机科学 2024-09-02 Ahmed Mustafa , Muhammad Tahir Rafique , Muhammad Ijlal Baig , Hasan Sajid , Muhammad Jawad Khan , Karam Dad Kallu

This article describes the results of a case study that applies Neural Network-based Optical Character Recognition (OCR) to scanned images of books printed between 1487 and 1870 by training the OCR engine OCRopus [@breuel2013high] on the…

计算与语言 · 计算机科学 2017-04-11 U. Springmann , A. Lüdeling

This paper presents a complete Optical Character Recognition (OCR) system for camera captured image/graphics embedded textual documents for handheld devices. At first, text regions are extracted and skew corrected. Then, these regions are…

计算机视觉与模式识别 · 计算机科学 2011-09-16 Ayatullah Faruk Mollah , Nabamita Majumder , Subhadip Basu , Mita Nasipuri

The Bavarian Academy of Sciences and Humanities aims to digitize its Medieval Latin Dictionary. This dictionary entails record cards referring to lemmas in medieval Latin, a low-resource language. A crucial step of the digitization process…

Arabic Optical Character Recognition (OCR) is essential for converting vast amounts of Arabic print media into digital formats. However, training modern OCR models, especially powerful vision-language models, is hampered by the lack of…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Omer Nacar , Yasser Al-Habashi , Serry Sibaee , Adel Ammar , Wadii Boulila

Ancient scripts, e.g., Egyptian hieroglyphs, Oracle Bone Inscriptions, and Ancient Greek inscriptions, serve as vital carriers of human civilization, embedding invaluable historical and cultural information. Automating ancient script image…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Xiaolei Diao , Rite Bo , Yanling Xiao , Lida Shi , Zhihan Zhou , Hao Xu , Chuntao Li , Xiongfeng Tang , Massimo Poesio , Cédric M. John , Daqian Shi