中文
相关论文

相关论文: Afro-MNIST: Synthetic generation of MNIST-style da…

200 篇论文

In this paper, we disseminate a new handwritten digits-dataset, termed Kannada-MNIST, for the Kannada script, that can potentially serve as a direct drop-in replacement for the original MNIST dataset. In addition to this dataset, we…

计算机视觉与模式识别 · 计算机科学 2019-08-06 Vinay Uday Prabhu

In this letter, we contribute a multi-language handwritten digit recognition dataset named MNIST-MIX, which is the largest dataset of the same type in terms of both languages and data samples. With the same data format with MNIST, MNIST-MIX…

计算机视觉与模式识别 · 计算机科学 2021-01-28 Weiwei Jiang

The MNIST dataset has become a standard benchmark for learning, classification and computer vision systems. Contributing to its widespread adoption are the understandable and intuitive nature of the task, its relatively small size and…

计算机视觉与模式识别 · 计算机科学 2017-03-02 Gregory Cohen , Saeed Afshar , Jonathan Tapson , André van Schaik

Neural networks are often benchmarked using standard datasets such as MNIST, FashionMNIST, or other variants of MNIST, which, while accessible, are limited to generic classes such as digits or clothing items. For researchers working on…

机器学习 · 计算机科学 2025-07-17 Pouya Shaeri , Arash Karimi , Ariane Middel

We present Typography-MNIST (TMNIST), a dataset comprising of 565,292 MNIST-style grayscale images representing 1,812 unique glyphs in varied styles of 1,355 Google-fonts. The glyph-list contains common characters from over 150 of the…

计算机视觉与模式识别 · 计算机科学 2022-02-17 Nimish Magre , Nicholas Brown

The advancement of speech technologies has been remarkable, yet its integration with African languages remains limited due to the scarcity of African speech corpora. To address this issue, we present AfroDigits, a minimalist,…

Training data for machine learning models can come from many different sources, which can be of dubious quality. For resource-rich languages like English, there is a lot of data available, so we can afford to throw out the dubious data. For…

计算与语言 · 计算机科学 2021-03-31 Andrew Zupon , Evan Crew , Sandy Ritchie

As low-resourced languages are increasingly incorporated into NLP research, there is an emphasis on collecting large-scale datasets. But in prioritizing quantity over quality, we risk 1) building language technologies that perform poorly…

Reproducible benchmarks are crucial in driving progress of machine translation research. However, existing machine translation benchmarks have been mostly limited to high-resource or well-represented languages. Despite an increasing…

计算与语言 · 计算机科学 2021-09-13 Machel Reid , Junjie Hu , Graham Neubig , Yutaka Matsuo

We present lightweight flow matching multilingual text-to-speech (TTS) systems for Ojibwe, Mi'kmaq, and Maliseet, three Indigenous languages in North America. Our results show that training a multilingual TTS model on three typologically…

计算与语言 · 计算机科学 2025-02-06 Shenran Wang , Changbing Yang , Mike Parkhill , Chad Quinn , Christopher Hammerly , Jian Zhu

Scaling multilingual representation learning beyond the hundred most frequent languages is challenging, in particular to cover the long tail of low-resource languages. A promising approach has been to train one-for-all multilingual models…

计算与语言 · 计算机科学 2022-05-26 Kevin Heffernan , Onur Çelebi , Holger Schwenk

Modern speech synthesis techniques can produce natural-sounding speech given sufficient high-quality data and compute resources. However, such data is not readily available for many languages. This paper focuses on speech synthesis for…

计算与语言 · 计算机科学 2022-07-05 Perez Ogayo , Graham Neubig , Alan W Black

Driven by advances in recording technology, large-scale high-dimensional datasets have emerged across many scientific disciplines. Especially in biology, clustering is often used to gain insights into the structure of such datasets, for…

机器学习 · 计算机科学 2024-10-22 Polina Turishcheva , Laura Hansel , Martin Ritzert , Marissa A. Weis , Alexander S. Ecker

Building effective neural machine translation (NMT) models for very low-resourced and morphologically rich African indigenous languages is an open challenge. Besides the issue of finding available resources for them, a lot of work is put…

计算与语言 · 计算机科学 2021-03-18 Bonaventure F. P. Dossou , Chris C. Emezue

Neural retrieval and GPT-style generative models rely on large, high-quality supervised data, which is still scarce for low-resource languages such as Amharic. We release an Amharic data resource consisting of two datasets that supports…

计算与语言 · 计算机科学 2026-02-11 Tilahun Yeshambel , Moncef Garouani , Josiane Mothe

We introduce the Oracle-MNIST dataset, comprising of 28$\times $28 grayscale images of 30,222 ancient characters from 10 categories, for benchmarking pattern classification, with particular challenges on image noise and distortion. The…

计算机视觉与模式识别 · 计算机科学 2024-02-14 Mei Wang , Weihong Deng

Stereotype repositories are critical to assess generative AI model safety, but currently lack adequate global coverage. It is imperative to prioritize targeted expansion, strategically addressing existing deficits, over merely increasing…

Automatic speech recognition (ASR) for African languages remains constrained by limited labeled data and the lack of systematic guidance on model selection, data scaling, and decoding strategies. Large pre-trained systems such as Whisper,…

In Africa, and the world at large, there is an increasing focus on developing Neural Machine Translation (NMT) systems to overcome language barriers. NMT for Low-resource language is particularly compelling as it involves learning with…

计算与语言 · 计算机科学 2023-08-28 Sakayo Toadoum Sari , Angela Fan , Lema Logamou Seknewna

Named Entity Recognition(NER) for low-resource languages aims to produce robust systems for languages where there is limited labeled training data available, and has been an area of increasing interest within NLP. Data augmentation for…

计算与语言 · 计算机科学 2026-02-16 Gaurav Kamath , Sowmya Vajjala
‹ 上一页 1 2 3 10 下一页 ›