English
Related papers

Related papers: Oracle-MNIST: a Dataset of Oracle Characters for B…

200 papers

We present Typography-MNIST (TMNIST), a dataset comprising of 565,292 MNIST-style grayscale images representing 1,812 unique glyphs in varied styles of 1,355 Google-fonts. The glyph-list contains common characters from over 150 of the…

Computer Vision and Pattern Recognition · Computer Science 2022-02-17 Nimish Magre , Nicholas Brown

Oracle bone inscriptions(OBI) is the earliest developed writing system in China, bearing invaluable written exemplifications of early Shang history and paleography. However, the task of deciphering OBI, in the current climate of the…

The MNIST dataset has become a standard benchmark for learning, classification and computer vision systems. Contributing to its widespread adoption are the understandable and intuitive nature of the task, its relatively small size and…

Computer Vision and Pattern Recognition · Computer Science 2017-03-02 Gregory Cohen , Saeed Afshar , Jonathan Tapson , André van Schaik

We report generation of a MNIST [4] compatible data set [1] for Tamil vowels to enable building a classification DNN or other such ML/AI deep learning [2] models for Tamil OCR/Handwriting applications. We report the capability of the 60,000…

Computer Vision and Pattern Recognition · Computer Science 2020-06-18 Muthiah Annamalai

Oracle bone script, one of the earliest known forms of ancient Chinese writing, presents invaluable research materials for scholars studying the humanities and geography of the Shang Dynasty, dating back 3,000 years. The immense historical…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Pengjie Wang , Kaile Zhang , Xinyu Wang , Shengwei Han , Yongge Liu , Jinpeng Wan , Haisu Guan , Zhebin Kuang , Lianwen Jin , Xiang Bai , Yuliang Liu

The research presents an overhead view of 10 important objects and follows the general formatting requirements of the most popular machine learning task: digit recognition with MNIST. This dataset offers a public benchmark extracted from…

Computer Vision and Pattern Recognition · Computer Science 2021-02-09 David Noever , Samantha E. Miller Noever

The short note presents an image classification dataset consisting of 10 executable code varieties and approximately 50,000 virus examples. The malicious classes include 9 families of computer viruses and one benign set. The image…

Cryptography and Security · Computer Science 2021-03-02 David Noever , Samantha E. Miller Noever

The Virus-MNIST data set is a collection of thumbnail images that is similar in style to the ubiquitous MNIST hand-written digits. These, however, are cast by reshaping possible malware code into an image array. Naturally, it is poised to…

Machine Learning · Computer Science 2021-11-04 Erik Larsen , Korey MacVittie , John Lilly

Quantum machine learning (QML) has emerged as a promising domain to leverage the computational capabilities of quantum systems to solve complex classification tasks. In this work, we present the first comprehensive QML study by benchmarking…

Quantum Physics · Physics 2025-03-21 Gurinder Singh , Hongni Jin , Kenneth M. Merz

We propose a new quantum neural network for image classification, which is able to classify the parity of the MNIST dataset with full resolution with a test accuracy of up to 97.5% without any classical pre-processing or post-processing.…

Quantum Physics · Physics 2025-05-22 Paolo Alessandro Xavier Tognini , Leonardo Banchi , Giacomo De Palma

Foundational to the Chinese language and culture, Chinese characters encompass extraordinarily extensive and ever-expanding categories, with the latest Chinese GB18030-2022 standard containing 87,887 categories. The accurate recognition of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Yuyi Zhang , Yongxin Shi , Peirong Zhang , Yixin Zhao , Zhenhua Yang , Lianwen Jin

Due to the sparsity of features, noise has proven to be a great inhibitor in the classification of handwritten characters. To combat this, most techniques perform denoising of the data before classification. In this paper, we consolidate…

Computer Vision and Pattern Recognition · Computer Science 2019-08-27 Qun Liu , Edward Collier , Supratik Mukhopadhyay

We introduce MedMNIST v2, a large-scale MNIST-like dataset collection of standardized biomedical images, including 12 datasets for 2D and 6 datasets for 3D. All images are pre-processed into a small size of 28x28 (2D) or 28x28x28 (3D) with…

Computer Vision and Pattern Recognition · Computer Science 2023-02-20 Jiancheng Yang , Rui Shi , Donglai Wei , Zequan Liu , Lin Zhao , Bilian Ke , Hanspeter Pfister , Bingbing Ni

In this letter, we contribute a multi-language handwritten digit recognition dataset named MNIST-MIX, which is the largest dataset of the same type in terms of both languages and data samples. With the same data format with MNIST, MNIST-MIX…

Computer Vision and Pattern Recognition · Computer Science 2021-01-28 Weiwei Jiang

Constructing historical language models (LMs) plays a crucial role in aiding archaeological provenance studies and understanding ancient cultures. However, existing resources present major challenges for training effective LMs on historical…

Computation and Language · Computer Science 2025-08-25 Xiaolei Diao , Zhihan Zhou , Lida Shi , Ting Wang , Ruihua Qi , Hao Xu , Daqian Shi

Recognizing handwritten digits is a challenging task primarily due to the diversity of writing styles and the presence of noisy images. The widely used MNIST dataset, which is commonly employed as a benchmark for this task, includes…

Computer Vision and Pattern Recognition · Computer Science 2023-07-28 Amarnath R , Vinay Kumar

We introduce OmniPrint, a synthetic data generator of isolated printed characters, geared toward machine learning research. It draws inspiration from famous datasets such as MNIST, SVHN and Omniglot, but offers the capability of generating…

Computer Vision and Pattern Recognition · Computer Science 2022-01-19 Haozhe Sun , Wei-Wei Tu , Isabelle Guyon

We present the Noisy Ostracods, a noisy dataset for genus and species classification of crustacean ostracods with specialists' annotations. Over the 71466 specimens collected, 5.58% of them are estimated to be noisy (possibly problematic)…

Machine Learning · Computer Science 2024-12-04 Jiamian Hu , Yuanyuan Hong , Yihua Chen , He Wang , Moriaki Yasuhara

The success of CNN-based architecture on image classification in learning and extracting features made them so popular these days, but the task of image classification becomes more challenging when we apply state of art models to classify…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Ashkan Ganj , Mohsen Ebadpour , Mahdi Darvish , Hamid Bahador

This technical report presents the 600K-KS-OCR Dataset, a large-scale synthetic corpus comprising approximately 602,000 word-level segmented images designed for training and evaluating optical character recognition systems targeting…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Haq Nawaz Malik
‹ Prev 1 2 3 10 Next ›