English
Related papers

Related papers: Oracle-MNIST: a Dataset of Oracle Characters for B…

200 papers

Neural networks are often benchmarked using standard datasets such as MNIST, FashionMNIST, or other variants of MNIST, which, while accessible, are limited to generic classes such as digits or clothing items. For researchers working on…

Machine Learning · Computer Science 2025-07-17 Pouya Shaeri , Arash Karimi , Ariane Middel

Scene Text Image Super-resolution (STISR) aims to recover high-resolution (HR) scene text images with visually pleasant and readable text content from the given low-resolution (LR) input. Most existing works focus on recovering English…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Jianqi Ma , Zhetong Liang , Wangmeng Xiang , Xi Yang , Lei Zhang

Although the popular MNIST dataset [LeCun et al., 1994] is derived from the NIST database [Grother and Hanaoka, 1995], the precise processing steps for this derivation have been lost to time. We propose a reconstruction that is accurate…

Machine Learning · Computer Science 2019-11-06 Chhavi Yadav , Léon Bottou

Multiple-instance learning is a subset of weakly supervised learning where labels are applied to sets of instances rather than the instances themselves. Under the standard assumption, a set is positive only there is if at least one instance…

Machine Learning · Computer Science 2021-05-05 Daniel Grahn

Many localized languages struggle to reap the benefits of recent advancements in character recognition systems due to the lack of substantial amount of labeled training data. This is due to the difficulty in generating large amounts of…

Computer Vision and Pattern Recognition · Computer Science 2020-08-10 Vinoj Jayasundara , Sandaru Jayasekara , Hirunima Jayasekara , Jathushan Rajasegaran , Suranga Seneviratne , Ranga Rodrigo

Oracle bone script is the earliest-known Chinese writing system of the Shang dynasty and is precious to archeology and philology. However, real-world scanned oracle data are rare and few experts are available for annotation which make the…

Computer Vision and Pattern Recognition · Computer Science 2022-05-16 Mei Wang , Weihong Deng , Cheng-Lin Liu

The MNIST dataset containing thousands of handwritten digit images is still a fundamental benchmark for evaluating various pattern-recognition and image-classification models. Linear separability is a key concept in many statistical and…

Machine Learning · Computer Science 2026-03-16 Ákos Hajnal

Twenty-three machine learning algorithms were trained then scored to establish baseline comparison metrics and to select an image classification algorithm worthy of embedding into mission-critical satellite imaging systems. The…

Computer Vision and Pattern Recognition · Computer Science 2021-10-22 Erik Larsen , David Noever , Korey MacVittie , John Lilly

Developing effective scene text detection and recognition models hinges on extensive training data, which can be both laborious and costly to obtain, especially for low-resourced languages. Conventional methods tailored for Latin characters…

Computer Vision and Pattern Recognition · Computer Science 2024-10-25 Vannkinh Nom , Souhail Bakkali , Muhammad Muzzamil Luqman , Mickaël Coustaty , Jean-Marc Ogier

Driven by advances in recording technology, large-scale high-dimensional datasets have emerged across many scientific disciplines. Especially in biology, clustering is often used to gain insights into the structure of such datasets, for…

Machine Learning · Computer Science 2024-10-22 Polina Turishcheva , Laura Hansel , Martin Ritzert , Marissa A. Weis , Alexander S. Ecker

In this paper, we disseminate a new handwritten digits-dataset, termed Kannada-MNIST, for the Kannada script, that can potentially serve as a direct drop-in replacement for the original MNIST dataset. In addition to this dataset, we…

Computer Vision and Pattern Recognition · Computer Science 2019-08-06 Vinay Uday Prabhu

An image dataset of 10 different size molecules, where each molecule has 2,000 structural variants, is generated from the 2D cross-sectional projection of Molecular Dynamics trajectories. The purpose of this dataset is to provide a…

Image and Video Processing · Electrical Eng. & Systems 2019-11-19 Yan Zhang , Steve Farrell , Michael Crowley , Lee Makowski , Jack Deslippe

An ongoing challenge in current natural language processing is how its major advancements tend to disproportionately favor resource-rich languages, leaving a significant number of under-resourced languages behind. Due to the lack of…

Computation and Language · Computer Science 2023-02-13 Ruoyu Xie , Antonios Anastasopoulos

Understanding humanity's earliest writing systems is crucial for reconstructing civilization's origins, yet many ancient scripts remain undeciphered. Oracle Bone Script (OBS) from China's Shang dynasty exemplifies this challenge: only…

Information Retrieval · Computer Science 2026-04-14 Yin Wu , Gangjian Zhang , Jiayu Chen , Chang Xu , Yuyu Luo , Nan Tang , Hui Xiong

Large-scale medical imaging datasets have accelerated deep learning (DL) for medical image analysis. However, the large scale of these datasets poses a challenge for researchers, resulting in increased storage and bandwidth requirements for…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Pranav Kulkarni , Adway Kanhere , Eliot Siegel , Paul H. Yi , Vishwa S. Parekh

Label noise remains a challenge for training robust classification models. Most methods for mitigating label noise have been benchmarked using primarily datasets with synthetic noise. While the need for datasets with realistic noise…

Computation and Language · Computer Science 2024-10-24 Alicja Rączkowska , Aleksandra Osowska-Kurczab , Jacek Szczerbiński , Kalina Jasinska-Kobus , Klaudia Nazarko

Many real-world applications involve the use of Optical Character Recognition (OCR) engines to transform handwritten images into transcripts on which downstream Natural Language Processing (NLP) models are applied. In this process, OCR…

Computation and Language · Computer Science 2021-07-16 Guowei Xu , Wenbiao Ding , Weiping Fu , Zhongqin Wu , Zitao Liu

Origami is becoming more and more relevant to research. However, there is no public dataset yet available and there hasn't been any research on this topic in machine learning. We constructed an origami dataset using images from the…

Computer Vision and Pattern Recognition · Computer Science 2021-01-15 Daniel Ma , Gerald Friedland , Mario Michael Krell

Oracle bone inscriptions (OBIs) are the earliest known form of Chinese characters and serve as a valuable resource for research in anthropology and archaeology. However, most excavated fragments are severely degraded due to thousands of…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Jinhao Li , Zijian Chen , Tingzhu Chen , Zhiji Liu , Changbo Wang

We present an approach to effectively use millions of images with noisy annotations in conjunction with a small subset of cleanly-annotated images to learn powerful image representations. One common approach to combine clean and noisy data…

Computer Vision and Pattern Recognition · Computer Science 2017-04-11 Andreas Veit , Neil Alldrin , Gal Chechik , Ivan Krasin , Abhinav Gupta , Serge Belongie