中文
相关论文

相关论文: An open dataset for oracle bone script recognition…

200 篇论文

Reading comprehension by machine has been widely studied, but machine comprehension of spoken content is still a less investigated problem. In this paper, we release Open-Domain Spoken Question Answering Dataset (ODSQA) with more than three…

计算与语言 · 计算机科学 2018-08-08 Chia-Hsuan Lee , Shang-Ming Wang , Huan-Cheng Chang , Hung-Yi Lee

Optical character recognition (OCR) for historical documents is a complex procedure subject to a unique set of material issues, including inconsistencies in typefaces and low quality scanning. Consequently, even the most sophisticated OCR…

计算与语言 · 计算机科学 2020-04-27 Alberto Poncelas , Mohammad Aboomar , Jan Buts , James Hadley , Andy Way

The development of accurate and efficient machine learning models for predicting the structure and properties of molecular crystals has been hindered by the scarcity of publicly available datasets of structures with property labels. To…

Optical Character Recognition (OCR) is an established task with the objective of identifying the text present in an image. While many off-the-shelf OCR models exist, they are often trained for either scientific (e.g., formulae) or generic…

计算与语言 · 计算机科学 2024-03-26 Nan Zhang , Connor Heaton , Sean Timothy Okonsky , Prasenjit Mitra , Hilal Ezgi Toraman

This paper presents a novel approach to Chinese characters through the lens of physics, network analysis, and natural systems. Computational analysis of over 6,000 characters identified 422 elemental characters as fundamental building…

物理与社会 · 物理学 2025-02-28 Wen G. Gong

Determining the chronology of ancient handwritten manuscripts is essential for reconstructing the evolution of ideas. For the Dead Sea Scrolls, this is particularly important. However, there is an almost complete lack of date-bearing…

Facial expression recognition (FER) is a challenging task due to pervasive occlusion and dataset biases. Especially when facial information is partially occluded, existing FER models struggle to extract effective facial features, leading to…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Huiyu Zhai , Xingxing Yang , Yalan Ye , Chenyang Li , Bin Fan , Changze Li

Nowadays document analysis and recognition remain challenging tasks. However, only a few datasets designed for text detection (TD) and optical character recognition (OCR) problems exist. In this paper we present Distorted Document Images…

计算机视觉与模式识别 · 计算机科学 2021-09-20 Ilia Zharikov , Filipp Nikitin , Ilia Vasiliev , Vladimir Dokholyan

Digitization of newspapers is of interest for many reasons including preservation of history, accessibility and search ability, etc. While digitization of documents such as scientific articles and magazines is prevalent in literature, one…

计算机视觉与模式识别 · 计算机科学 2022-02-04 Wenzhen Zhu , Negin Sokhandan , Guang Yang , Sujitha Martin , Suchitra Sathyanarayana

Synthetic image generation has recently experienced significant improvements in domains such as natural image or art generation. However, the problem of figure and diagram generation remains unexplored. A challenging aspect of generating…

计算机视觉与模式识别 · 计算机科学 2022-10-26 Juan A. Rodriguez , David Vazquez , Issam Laradji , Marco Pedersoli , Pau Rodriguez

We present the largest publicly available synthetic OCR benchmark dataset for Indic languages. The collection contains a total of 90k images and their ground truth for 23 Indic languages. OCR model validation in Indic languages require a…

计算机视觉与模式识别 · 计算机科学 2022-05-06 Naresh Saini , Promodh Pinto , Aravinth Bheemaraj , Deepak Kumar , Dhiraj Daga , Saurabh Yadav , Srihari Nagaraj

Most existing online writer-identification systems require that the text content is supplied in advance and rely on separately designed features and classifiers. The identifications are based on lines of text, entire paragraphs, or entire…

计算机视觉与模式识别 · 计算机科学 2015-05-20 Weixin Yang , Lianwen Jin , Manfei Liu

Traditional OCR systems (OCR-1.0) are increasingly unable to meet people's usage due to the growing demand for intelligent processing of man-made optical characters. In this paper, we collectively refer to all artificial optical signals…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Haoran Wei , Chenglong Liu , Jinyue Chen , Jia Wang , Lingyu Kong , Yanming Xu , Zheng Ge , Liang Zhao , Jianjian Sun , Yuang Peng , Chunrui Han , Xiangyu Zhang

At present, recognition of the Bangla handwriting compound character has been an essential issue for many years. In recent years there have been application-based researches in machine learning, and deep learning, which is gained interest,…

计算机视觉与模式识别 · 计算机科学 2024-12-05 Jannatul Ferdous , Suvrajit Karmaker , A K M Shahariar Azad Rabby , Syed Akhter Hossain

Oscar Wilde said, "The difference between literature and journalism is that journalism is unreadable, and literature is not read." Unfortunately, The digitally archived journalism of Oscar Wilde's 19th century often has no or poor quality…

计算与语言 · 计算机科学 2025-02-24 Jonathan Bourne

Twenty-five hundred years ago, the paperwork of the Achaemenid Empire was recorded on clay tablets. In 1933, archaeologists from the University of Chicago's Oriental Institute (OI) found tens of thousands of these tablets and fragments…

计算机视觉与模式识别 · 计算机科学 2025-02-04 Edward C. Williams , Grace Su , Sandra R. Schloen , Miller C. Prosser , Susanne Paulus , Sanjay Krishnan

Face recognition has recently become ubiquitous in many scenes for authentication or security purposes. Meanwhile, there are increasing concerns about the privacy of face images, which are sensitive biometric data that should be carefully…

密码学与安全 · 计算机科学 2022-01-31 Qi Zhao , Huanhao Li , Zhipeng Yu , Chi Man Woo , Tianting Zhong , Shengfu Cheng , Yuanjin Zheng , Honglin Liu , Jie Tian , Puxiang Lai

In medieval India, the Marathi language was written using the Modi script. The texts written in Modi script include extensive knowledge about medieval sciences, medicines, land records and authentic evidence about Indian history. Around 40…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Harshal Kausadikar , Tanvi Kale , Onkar Susladkar , Sparsh Mittal

Online Arabic cursive character recognition is still a big challenge due to the existing complexities including Arabic cursive script styles, writing speed, writer mood and so forth. Due to these unavoidable constraints, the accuracy of…

计算机视觉与模式识别 · 计算机科学 2021-01-18 Amjad Rehman

All fields of knowledge are being impacted by Artificial Intelligence. In particular, the Deep Learning paradigm enables the development of data analysis tools that support subject matter experts in a variety of sectors, from physics up to…