English

Afro-MNIST: Synthetic generation of MNIST-style datasets for low-resource languages

Computer Vision and Pattern Recognition 2020-09-29 v1 Machine Learning

Abstract

We present Afro-MNIST, a set of synthetic MNIST-style datasets for four orthographies used in Afro-Asiatic and Niger-Congo languages: Ge`ez (Ethiopic), Vai, Osmanya, and N'Ko. These datasets serve as "drop-in" replacements for MNIST. We also describe and open-source a method for synthetic MNIST-style dataset generation from single examples of each digit. These datasets can be found at https://github.com/Daniel-Wu/AfroMNIST. We hope that MNIST-style datasets will be developed for other numeral systems, and that these datasets vitalize machine learning education in underrepresented nations in the research community.

Keywords

Cite

@article{arxiv.2009.13509,
  title  = {Afro-MNIST: Synthetic generation of MNIST-style datasets for low-resource languages},
  author = {Daniel J Wu and Andrew C Yang and Vinay U Prabhu},
  journal= {arXiv preprint arXiv:2009.13509},
  year   = {2020}
}

Comments

10 pages, 11 figures, presented as a workshop paper at Practical Machine Learning for Developing Countries @ ICLR 2020