中文
相关论文

相关论文: An open dataset for oracle bone script recognition…

200 篇论文

Multilingual OCR and information extraction from receipts remains challenging, particularly for complex scripts like Arabic. We introduce \dataset, a comprehensive dataset designed for Arabic-English receipt understanding comprising 20,000…

Machine Reading Comprehension (MRC) has become enormously popular recently and has attracted a lot of attention. However, the existing reading comprehension datasets are mostly in English. In this paper, we introduce a Span-Extraction…

计算与语言 · 计算机科学 2019-11-05 Yiming Cui , Ting Liu , Wanxiang Che , Li Xiao , Zhipeng Chen , Wentao Ma , Shijin Wang , Guoping Hu

Scholars in the humanities rely heavily on ancient manuscripts to study history, religion, and socio-political structures in the past. Many efforts have been devoted to digitizing these precious manuscripts using OCR technology, but most…

计算与语言 · 计算机科学 2026-05-19 Queenie Luo , Yung-Sung Chuang

The recognition of cursive script is regarded as a subtle task in optical character recognition due to its varied representation. Every cursive script has different nature and associated challenges. As Urdu is one of cursive language that…

计算机视觉与模式识别 · 计算机科学 2017-05-17 Saad Bin Ahmed , Saeeda Naz , Salahuddin Swati , Muhammad Imran Razzak

Chinese Spelling Correction (CSC) is gaining increasing attention due to its promise of automatically detecting and correcting spelling errors in Chinese texts. Despite its extensive use in many applications, like search engines and optical…

计算与语言 · 计算机科学 2022-10-24 Wangjie Jiang , Zhihao Ye , Zijing Ou , Ruihui Zhao , Jianguang Zheng , Yi Liu , Siheng Li , Bang Liu , Yujiu Yang , Yefeng Zheng

Handwriting recognition is of crucial importance to both Human Computer Interaction (HCI) and paperwork digitization. In the general field of Optical Character Recognition (OCR), handwritten Chinese character recognition faces tremendous…

计算机视觉与模式识别 · 计算机科学 2021-03-18 Boxiang Dong , Aparna S. Varde , Danilo Stevanovic , Jiayin Wang , Liang Zhao

Computational replication of Chinese calligraphy remains challenging. Existing methods falter, either creating high-quality isolated characters while ignoring page-level aesthetics like ligatures and spacing, or attempting page synthesis at…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Tianshuo Xu , Kai Wang , Zhifei Chen , Leyi Wu , Tianshui Wen , Fei Chao , Ying-Cong Chen

Handwritten Digit Recognition (HDR) is one of the most challenging tasks in the domain of Optical Character Recognition (OCR). Irrespective of language, there are some inherent challenges of HDR, which mostly arise due to the variations in…

Origami is becoming more and more relevant to research. However, there is no public dataset yet available and there hasn't been any research on this topic in machine learning. We constructed an origami dataset using images from the…

计算机视觉与模式识别 · 计算机科学 2021-01-15 Daniel Ma , Gerald Friedland , Mario Michael Krell

Optical character recognition (OCR) is a process of converting analogue documents into digital using document images. Currently, many commercial and non-commercial OCR systems exist for both handwritten and printed copies for different…

计算机视觉与模式识别 · 计算机科学 2021-05-11 Farisa Benta Safir , Abu Quwsar Ohi , M. F. Mridha , Muhammad Mostafa Monowar , Md. Abdul Hamid

Stroke extraction of Chinese characters plays an important role in the field of character recognition and generation. The most existing character stroke extraction methods focus on image morphological features. These methods usually lead to…

计算机视觉与模式识别 · 计算机科学 2023-07-11 Meng Li , Yahan Yu , Yi Yang , Guanghao Ren , Jian Wang

The history of the Korean language is characterized by a discrepancy between its spoken and written forms and a pivotal shift from Chinese characters to the Hangul alphabet. However, this linguistic evolution has remained largely unexplored…

计算与语言 · 计算机科学 2026-05-04 Seyoung Song , Nawon Kim , Songeun Chae , Kiwoong Park , Jiho Jin , Haneul Yoo , Kyunghyun Cho , Alice Oh

The objective of the paper is to recognize handwritten samples of Roman numerals using Tesseract open source Optical Character Recognition (OCR) engine. Tesseract is trained with data samples of different persons to generate one…

计算机视觉与模式识别 · 计算机科学 2010-03-31 Sandip Rakshit , Amitava Kundu , Mrinmoy Maity , Subhajit Mandal , Satwika Sarkar , Subhadip Basu

Bangla handwriting recognition is becoming a very important issue nowadays. It is potentially a very important task specially for Bangla speaking population of Bangladesh and West Bengal. By keeping that in our mind we are introducing a…

计算与语言 · 计算机科学 2017-04-03 Mithun Biswas , Rafiqul Islam , Gautam Kumar Shom , Md Shopon , Nabeel Mohammed , Sifat Momen , Md Anowarul Abedin

Optical character recognition (OCR) is crucial for a deeper access to historical collections. OCR needs to account for orthographic variations, typefaces, or language evolution (i.e., new letters, word spellings), as the main source of…

计算与语言 · 计算机科学 2021-02-02 Lijun Lyu , Maria Koutraki , Martin Krickl , Besnik Fetahu

Recently, great success has been achieved in offline handwritten Chinese character recognition by using deep learning methods. Chinese characters are mainly logographic and consist of basic radicals, however, previous research mostly…

计算机视觉与模式识别 · 计算机科学 2018-08-14 Wenchao Wang , Jianshu Zhang , Jun Du , Zi-Rui Wang , Yixing Zhu

In the present work, we have used Tesseract 2.01 open source Optical Character Recognition (OCR) Engine under Apache License 2.0 for recognition of handwriting samples of lower case Roman script. Handwritten isolated and free-flow text…

计算机视觉与模式识别 · 计算机科学 2010-03-31 Sandip Rakshit , Subhadip Basu

In this paper we evaluate Optical Character Recognition (OCR) of 19th century Fraktur scripts without book-specific training using mixed models, i.e. models trained to recognize a variety of fonts and typesets from previously unseen…

计算机视觉与模式识别 · 计算机科学 2018-10-09 Christian Reul , Uwe Springmann , Christoph Wick , Frank Puppe

We present a novel approach to OCR(Optical Character Recognition) of Korean character, Hangul. As a phonogram, Hangul can represent 11,172 different characters with only 52 graphemes, by describing each character with a combination of the…

计算机视觉与模式识别 · 计算机科学 2022-09-29 Geonuk Kim , Jaemin Son , Kanghyu Lee , Jaesik Min

Bronze inscriptions (BI), engraved on ritual vessels, constitute a crucial stage of early Chinese writing and provide indispensable evidence for archaeological and historical studies. However, automatic BI recognition remains difficult due…

计算机视觉与模式识别 · 计算机科学 2025-10-03 Rixin Zhou , Peiqiang Qiu , Qian Zhang , Chuntao Li , Xi Yang