中文
相关论文

相关论文: Ancient Script Image Recognition and Processing: A…

200 篇论文

Standard natural language processing (NLP) pipelines operate on symbolic representations of language, which typically consist of sequences of discrete tokens. However, creating an analogous representation for ancient logographic writing…

计算与语言 · 计算机科学 2026-01-29 Danlu Chen , Freda Shi , Aditi Agarwal , Jacobo Myerston , Taylor Berg-Kirkpatrick

With the rapid increase of transnational communication and cooperation, people frequently encounter multilingual scenarios in various situations. In this paper, we are concerned with a relatively new problem: script identification at word…

计算机视觉与模式识别 · 计算机科学 2015-05-13 Baoguang Shi , Cong Yao , Chengquan Zhang , Xiaowei Guo , Feiyue Huang , Xiang Bai

This paper focuses on the problem of script identification in unconstrained scenarios. Script identification is an important prerequisite to recognition, and an indispensable condition for automatic text understanding systems designed for…

计算机视觉与模式识别 · 计算机科学 2016-02-25 Lluis Gomez , Dimosthenis Karatzas

Oracle bone script, one of the earliest known forms of ancient Chinese writing, presents invaluable research materials for scholars studying the humanities and geography of the Shang Dynasty, dating back 3,000 years. The immense historical…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Pengjie Wang , Kaile Zhang , Xinyu Wang , Shengwei Han , Yongge Liu , Jinpeng Wan , Haisu Guan , Zhebin Kuang , Lianwen Jin , Xiang Bai , Yuliang Liu

Ancient script images often suffer from severe background noise, low contrast, and degradation caused by aging and environmental effects. In many cases, the foreground text and background exhibit similar visual characteristics, making the…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Bapu D. Chendage , Rajivkumar S. Mente

Encoded (or ciphered) manuscripts are a special type of historical documents that contain encrypted text. The automatic recognition of this kind of documents is challenging because: 1) the cipher alphabet changes from one document to…

计算机视觉与模式识别 · 计算机科学 2020-09-29 Mohamed Ali Souibgui , Alicia Fornés , Yousri Kessentini , Crina Tudor

Historical Document Processing is the process of digitizing written material from the past for future use by historians and other scholars. It incorporates algorithms and software tools from various subfields of computer science, including…

计算机视觉与模式识别 · 计算机科学 2020-09-14 James P. Philips , Nasseh Tabrizi

Oracle bone inscriptions (OBIs) contain some of the oldest characters in the world and were used in China about 3000 years ago. As an ancient form of literature, OBIs store a lot of information that can help us understand the world history,…

计算机视觉与模式识别 · 计算机科学 2021-05-05 Yoshiyuki Fujikawa , Hengyi Li , Xuebin Yue , Aravinda C , Amar Prabhu G , Lin Meng

The OpenITI team has achieved Optical Character Recognition (OCR) accuracy rates for classical Arabic-script texts in the high nineties. These numbers are based on our tests of seven different Arabic-script texts of varying quality and…

计算机视觉与模式识别 · 计算机科学 2017-03-29 Maxim Romanov , Matthew Thomas Miller , Sarah Bowen Savant , Benjamin Kiessling

The history of text can be traced back over thousands of years. Rich and precise semantic information carried by text is important in a wide range of vision-based application scenarios. Therefore, text recognition in natural scenes has been…

计算机视觉与模式识别 · 计算机科学 2020-12-04 Xiaoxue Chen , Lianwen Jin , Yuanzhi Zhu , Canjie Luo , Tianwei Wang

Exploring and understanding efficient image representations is a long-standing challenge in computer vision. While deep learning has achieved remarkable progress across image understanding tasks, its internal representations are often…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Chenyuan Qu , Hao Chen , Jianbo Jiao

In this paper, we present an Optical Character Recognition (OCR) system specifically designed for the accurate recognition and digitization of Greek polytonic texts. By leveraging the combined strengths of convolutional layers for feature…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Perifanos Konstantinos , Goutsos Dionisis

As one of the earliest ancient languages, Oracle Bone Script (OBS) encapsulates the cultural records and intellectual expressions of ancient civilizations. Despite the discovery of approximately 4,500 OBS characters, only about 1,600 have…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Caoshuo Li , Zengmao Ding , Xiaobin Hu , Bang Li , Donghao Luo , AndyPian Wu , Chaoyang Wang , Chengjie Wang , Taisong Jin , SevenShu , Yunsheng Wu , Yongge Liu , Rongrong Ji

The application of handwritten text recognition to historical works is highly dependant on accurate text line retrieval. A number of systems utilizing a robust baseline detection paradigm have emerged recently but the advancement of layout…

计算机视觉与模式识别 · 计算机科学 2019-07-10 Benjamin Kiessling , Daniel Stökl Ben Ezra , Matthew Thomas Miller

Constructing historical language models (LMs) plays a crucial role in aiding archaeological provenance studies and understanding ancient cultures. However, existing resources present major challenges for training effective LMs on historical…

计算与语言 · 计算机科学 2025-08-25 Xiaolei Diao , Zhihan Zhou , Lida Shi , Ting Wang , Ruihua Qi , Hao Xu , Daqian Shi

Humans readily recognize objects from sparse line drawings, a capacity that appears early in development and persists across cultures, suggesting neural rather than purely learned origins. Yet the computational mechanism by which the brain…

人工智能 · 计算机科学 2026-04-16 Seowung Leem , Lin Gu , Ruogu Fang

This report explores the latest advances in the field of digital document recognition. With the focus on printed document imagery, we discuss the major developments in optical character recognition (OCR) and document image…

计算机视觉与模式识别 · 计算机科学 2014-12-16 Eugene Borovikov

This paper presents our methodology and findings from three tasks across Optical Character Recognition (OCR) and Document Layout Analysis using advanced deep learning techniques. First, for the historical Hebrew fragments of the Dead Sea…

计算机视觉与模式识别 · 计算机科学 2025-08-15 Hylke Westerdijk , Ben Blankenborg , Khondoker Ittehadul Islam

Building robust recognizers for Arabic has always been challenging. We demonstrate the effectiveness of an end-to-end trainable CNN-RNN hybrid architecture in recognizing Arabic text in videos and natural scenes. We outperform previous…

计算机视觉与模式识别 · 计算机科学 2017-11-08 Mohit Jain , Minesh Mathew , C. V. Jawahar

While natural language understanding of long-form documents is still an open challenge, such documents often contain structural information that can inform the design of models for encoding them. Movie scripts are an example of such richly…

计算与语言 · 计算机科学 2020-05-01 Gayatri Bhat , Avneesh Saluja , Melody Dye , Jan Florjanczyk