中文
相关论文

相关论文: Ancient but Digitized: Developing Handwritten Opti…

200 篇论文

Handwritten Arabic manuscripts preserve the Arab world's intellectual and cultural heritage, and writer identification supports provenance, authenticity verification, and historical analysis. Using the Muharaf dataset of historical Arabic…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Hamza A. Abushahla , Ariel Justine N. Panopio , Layth Al-Khairulla , Mohamed I. AlHajri

Optical Character Recognition (OCR) is crucial to the National Library of Norway's (NLN) digitisation process as it converts scanned documents into machine-readable text. However, for the S\'ami documents in NLN's collection, the OCR…

计算与语言 · 计算机科学 2025-01-14 Tita Enstad , Trond Trosterud , Marie Iversdatter Røsok , Yngvil Beyer , Marie Roald

Handwritten text recognition (HTR) for Arabic-script languages still lags behind Latin-script HTR, despite recent advances in model architectures, datasets, and benchmarks. We show that data quality is a significant limiting factor in many…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Sana Al-azzawi , Elisa Barney , Marcus Liwicki

In handwritten character recognition, benchmark database plays an important role in evaluating the performance of various algorithms and the results obtained by various researchers. In Devnagari script, there is lack of such official…

计算机视觉与模式识别 · 计算机科学 2013-09-23 Vikas J. Dongre , Vijay H. Mankar

Handwriting Recognition enables a person to scribble something on a piece of paper and then convert it into text. If we look into the practical reality there are enumerable styles in which a character may be written. These styles can be…

计算机视觉与模式识别 · 计算机科学 2010-04-20 Rahul Kala , Harsh Vazirani , Anupam Shukla , Ritu Tiwari

An end-to-end, segmentation-free, deep learning model trained from scratch is proposed, leveraging DCNN for feature extraction, alongside Bidirectional Long-Short Term Memory (BLSTM) for sequence recognition and Connectionist Temporal…

计算机视觉与模式识别 · 计算机科学 2024-06-26 Sondos Aabed , Ahmad Khairaldin

Optical character recognition (OCR) has advanced rapidly with deep learning and multimodal models, yet most methods focus on well-resourced scripts such as Latin and Chinese. Ethnic minority languages remain underexplored due to complex…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Bonan Liu , Zeyu Zhang , Bingbing Meng , Han Wang , Hanshuo Zhang , Chengping Wang , Daji Ergu , Ying Cai

Script identification plays a vital role in applications that involve handwriting and document analysis within a multi-script and multi-lingual environment. Moreover, it exhibits a profound connection with human cognition. This paper…

计算机视觉与模式识别 · 计算机科学 2024-05-30 Miguel A. Ferrer , Abhijit Das , Moises Diaz , Aythami Morales , Cristina Carmona-Duarte , Umapada Pal

Despite increasing interest in Syriac studies and growing digital availability of Syriac texts, there is currently no up-to-date infrastructure for discovering, identifying, classifying, and referencing works of Syriac literature. The…

数字图书馆 · 计算机科学 2021-07-01 Nathan P. Gibson , David A. Michelson , Daniel L. Schwartz

Dysgraphia is a learning disorder that affects handwriting abilities, making it challenging for children to write legibly and consistently. Early detection and monitoring are crucial for providing timely support and interventions. This…

计算机视觉与模式识别 · 计算机科学 2024-11-22 Vydeki D , Divyansh Bhandari , Pranav Pratap Patil , Aarush Anand Kulkarni

Sumerian transliteration is a conventional system for representing a scholar's interpretation of a tablet in the Latin script. Thanks to visionary digital Assyriology projects such as ETCSL, CDLI, and Oracc, a large number of Sumerian…

计算与语言 · 计算机科学 2026-02-26 Cole Simmons , Richard Diehl Martinez , Dan Jurafsky

Arabic is a semitic language characterized by a complex and rich morphology. The exceptional degree of ambiguity in the writing system, the rich morphology, and the highly complex word formation process of roots and patterns all contribute…

计算机视觉与模式识别 · 计算机科学 2014-12-25 Ibrahim Abdelaziz , Sherif Abdou

Over the past few decades, large archives of paper-based documents such as books and newspapers have been digitized using Optical Character Recognition. This technology is error-prone, especially for historical documents. To correct OCR…

计算与语言 · 计算机科学 2023-08-01 Omri Suissa , Avshalom Elmalech , Maayan Zhitomirsky-Geffet

The recognition of unconstrained handwriting continues to be a difficult task for computers despite active research for several decades. This is because handwritten text offers great challenges such as character and word segmentation,…

神经与进化计算 · 计算机科学 2013-01-22 Yusuf Perwej

In this paper, we introduce the first phase of a new dataset for offline Arabic handwriting recognition. The aim is to collect a very large dataset of isolated Arabic words that covers all letters of the alphabet in all possible shapes…

计算机视觉与模式识别 · 计算机科学 2014-11-19 Mohamed E. Hussein , Marwan Torki , Ahmed Elsallamy , Mahmoud Fayyaz

Multilingual OCR and information extraction from receipts remains challenging, particularly for complex scripts like Arabic. We introduce \dataset, a comprehensive dataset designed for Arabic-English receipt understanding comprising 20,000…

Digital humanities scholars increasingly use Large Language Models for historical document digitization, yet lack appropriate evaluation frameworks for LLM-based OCR. Traditional metrics fail to capture temporal biases and period-specific…

计算机视觉与模式识别 · 计算机科学 2025-10-09 Maria Levchenko

We aim to investigate the performance of current OCR systems on low resource languages and low resource scripts. We introduce and make publicly available a novel benchmark, OCR4MT, consisting of real and synthetic data, enriched with noise,…

计算与语言 · 计算机科学 2022-03-15 Oana Ignat , Jean Maillard , Vishrav Chaudhary , Francisco Guzmán

This paper explores the application of synthetic data in the post-OCR domain on multiple fronts by conducting experiments to assess the impact of data volume, augmentation, and synthetic data generation methods on model performance.…

计算与语言 · 计算机科学 2024-08-14 Shuhao Guan , Derek Greene

In this paper, the problem of handwritten digit recognition has been addressed. However, the underlying language is Persian/Arabic, and the system with which this task is a capsule network (CapsNet) has recently emerged as a more advanced…

计算机视觉与模式识别 · 计算机科学 2019-12-20 Ali Ghofrani , Rahil Mahdian Toroghi