English
Related papers

Related papers: An open dataset for oracle bone script recognition…

200 papers

Thousands of users consult digital archives daily, but the information they can access is unrepresentative of the diversity of documentary history. The sequence-to-sequence architecture typically used for optical character recognition (OCR)…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Jacob Carlson , Tom Bryan , Melissa Dell

We present the Manuscripts of Handwritten Arabic~(Muharaf) dataset, which is a machine learning dataset consisting of more than 1,600 historic handwritten page images transcribed by experts in archival Arabic. Each document image is…

Computer Vision and Pattern Recognition · Computer Science 2025-02-06 Mehreen Saeed , Adrian Chan , Anupam Mijar , Joseph Moukarzel , Georges Habchi , Carlos Younes , Amin Elias , Chau-Wai Wong , Akram Khater

Telugu is a Dravidian language spoken by more than 80 million people worldwide. The optical character recognition (OCR) of the Telugu script has wide ranging applications including education, health-care, administration etc. The beautiful…

Computer Vision and Pattern Recognition · Computer Science 2018-12-27 Chandra Prakash Konkimalla , Manikanta Srikar Yellapragada , Trishal Gayam , Souraj Mandal , Sumohana S. Channappayya

Dongba pictographic is the only pictographic script still in use in the world. Its pictorial ideographic features carry rich cultural and contextual information. However, due to the lack of relevant datasets, research on semantic…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Xiaojun Bi , Shuo Li , Junyao Xing , Ziyue Wang , Fuwen Luo , Weizheng Qiao , Lu Han , Ziwei Sun , Peng Li , Yang Liu

Despite being proposed as early as 1959, COBOL (Common Business-Oriented Language) still predominantly acts as an integral part of the majority of operations of several financial, banking, and governmental organizations. To support the…

Software Engineering · Computer Science 2023-06-09 Mir Sameed Ali , Nikhil Manjunath , Sridhar Chimalakonda

Medical visual question answering (Med-VQA) has tremendous potential in healthcare. However, the development of this technology is hindered by the lacking of publicly-available and high-quality labeled datasets for training and evaluation.…

Computer Vision and Pattern Recognition · Computer Science 2021-02-19 Bo Liu , Li-Ming Zhan , Li Xu , Lin Ma , Yan Yang , Xiao-Ming Wu

The automatic recognition of tabular data in document images presents a significant challenge due to the diverse range of table styles and complex structures. Tables offer valuable content representation, enhancing the predictive…

Computer Vision and Pattern Recognition · Computer Science 2024-04-22 Avinash Anand , Raj Jaiswal , Pijush Bhuyan , Mohit Gupta , Siddhesh Bangar , Md. Modassir Imam , Rajiv Ratn Shah , Shin'ichi Satoh

We present OCR-Quality, a comprehensive human-annotated dataset designed for evaluating and developing OCR quality assessment methods. The dataset consists of 1,000 PDF pages converted to PNG images at 300 DPI, sampled from diverse…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Yulong Zhang

Previous works on font generation mainly focus on the standard print fonts where character's shape is stable and strokes are clearly separated. There is rare research on brush handwriting font generation, which involves holistic structure…

Computer Vision and Pattern Recognition · Computer Science 2022-04-25 Shaozu Yuan , Ruixue Liu , Meng Chen , Baoyang Chen , Zhijie Qiu , Xiaodong He

Scientific literature serves as a high-quality corpus, supporting a lot of Natural Language Processing (NLP) research. However, existing datasets are centered around the English language, which restricts the development of Chinese…

Computation and Language · Computer Science 2022-09-13 Yudong Li , Yuqing Zhang , Zhe Zhao , Linlin Shen , Weijie Liu , Weiquan Mao , Hui Zhang

Over the past few decades, large archives of paper-based historical documents, such as books and newspapers, have been digitized using the Optical Character Recognition (OCR) technology. Unfortunately, this broadly used technology is…

Computation and Language · Computer Science 2023-08-01 Omri Suissa , Maayan Zhitomirsky-Geffet , Avshalom Elmalech

Digitization of historical documents is a challenging task in many digital humanities projects. A popular approach for digitization is to scan the documents into images, and then convert images into text using Optical Character Recognition…

Human-Computer Interaction · Computer Science 2023-08-01 Omri Suissa , Avshalom Elmalech , Maayan Zhitomirsky-Geffet

We present Native Chinese Reader (NCR), a new machine reading comprehension (MRC) dataset with particularly long articles in both modern and classical Chinese. NCR is collected from the exam questions for the Chinese course in China's high…

Computation and Language · Computer Science 2021-12-15 Shusheng Xu , Yichen Liu , Xiaoyu Yi , Siyuan Zhou , Huizi Li , Yi Wu

Chinese ancient documents, invaluable carriers of millennia of Chinese history and culture, hold rich knowledge across diverse fields but face challenges in digitization and understanding, i.e., traditional methods only scan images, while…

Historical materials are abundant. Yet, piecing together how human knowledge has evolved and spread both diachronically and synchronically remains a challenge that can so far only be very selectively addressed. The vast volume of materials…

Machine Learning · Computer Science 2025-01-16 Oliver Eberle , Jochen Büttner , Hassan El-Hajj , Grégoire Montavon , Klaus-Robert Müller , Matteo Valleriani

The recognition of Chinese characters has always been a challenging task due to their huge variety and complex structures. The latest research proves that such an enormous character set can be decomposed into a collection of about 500…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Shaowei Wang , Guanjie Huang , Xiangyu Luo

The task of Chinese Spelling Check (CSC) is aiming to detect and correct spelling errors that can be found in the text. While manually annotating a high-quality dataset is expensive and time-consuming, thus the scale of the training dataset…

Computation and Language · Computer Science 2022-09-16 Piji Li

This paper describes the COCO-Text dataset. In recent years large-scale datasets like SUN and Imagenet drove the advancement of scene understanding and object recognition. The goal of COCO-Text is to advance state-of-the-art in text…

Computer Vision and Pattern Recognition · Computer Science 2016-06-21 Andreas Veit , Tomas Matera , Lukas Neumann , Jiri Matas , Serge Belongie

Bronze inscriptions from early China are fragmentary and difficult to date. We introduce BIRD(Bronze Inscription Restoration and Dating), a fully encoded dataset grounded in standard scholarly transcriptions and chronological labels. We…

Computation and Language · Computer Science 2025-11-04 Wenjie Hua , Hoang H. Nguyen , Gangyan Ge

Encoded (or ciphered) manuscripts are a special type of historical documents that contain encrypted text. The automatic recognition of this kind of documents is challenging because: 1) the cipher alphabet changes from one document to…

Computer Vision and Pattern Recognition · Computer Science 2020-09-29 Mohamed Ali Souibgui , Alicia Fornés , Yousri Kessentini , Crina Tudor