English
Related papers

Related papers: The first large scale collection of diverse Hausa …

200 papers

Despite the progress we have recorded in the last few years in multilingual natural language processing, evaluation is typically limited to a small set of languages with available datasets which excludes a large number of low-resource…

Computation and Language · Computer Science 2024-03-08 David Ifeoluwa Adelani , Hannah Liu , Xiaoyu Shen , Nikita Vassilyev , Jesujoba O. Alabi , Yanke Mao , Haonan Gao , Annie En-Shiun Lee

We present our journey in training a speech language model for Wolof, an underrepresented language spoken in West Africa, and share key insights. We first emphasize the importance of collecting large-scale, spontaneous, high-quality…

Computation and Language · Computer Science 2025-09-26 Yaya Sy , Dioula Doucouré , Christophe Cerisara , Irina Illina

Natural Language Processing (NLP) research has made great advancements in recent years with major breakthroughs that have established new benchmarks. However, these advances have mainly benefited a certain group of languages commonly…

Computation and Language · Computer Science 2023-05-02 Derguene Mbaye , Moussa Diallo , Thierno Ibrahima Diop

Through this paper, we seek to reduce the communication barrier between the hearing-impaired community and the larger society who are usually not familiar with sign language in the sub-Saharan region of Africa with the largest occurrences…

Computer Vision and Pattern Recognition · Computer Science 2022-08-16 Steven Kolawole , Opeyemi Osakuade , Nayan Saxena , Babatunde Kazeem Olorisade

Human biases have been shown to influence the performance of models and algorithms in various fields, including Natural Language Processing. While the study of this phenomenon is garnering focus in recent years, the available resources are…

Computation and Language · Computer Science 2024-08-15 Ana Sofia Evans , Helena Moniz , Luísa Coheur

We present the Thiomi Dataset, a large-scale multimodal corpus spanning ten African languages across four language families: Swahili, Kikuyu, Kamba, Kimeru, Luo, Maasai, Kipsigis, Somali (East Africa); Wolof (West Africa); and Fulani…

Computation and Language · Computer Science 2026-04-07 Hillary Mutisya , John Mugane , Gavin Nyamboga , Brian Chege , Maryruth Gathoni

African languages are numerous, complex and low-resourced. The datasets required for machine translation are difficult to discover, and existing research is hard to reproduce. Minimal attention has been given to machine translation for…

Computation and Language · Computer Science 2019-06-17 Laura Martinus , Jade Z. Abbott

Modern speech synthesis techniques can produce natural-sounding speech given sufficient high-quality data and compute resources. However, such data is not readily available for many languages. This paper focuses on speech synthesis for…

Computation and Language · Computer Science 2022-07-05 Perez Ogayo , Graham Neubig , Alan W Black

The advancement of speech technologies has been remarkable, yet its integration with African languages remains limited due to the scarcity of African speech corpora. To address this issue, we present AfroDigits, a minimalist,…

The hospitality industry in the Arab world increasingly relies on customer feedback to shape services, driving the need for advanced Arabic sentiment analysis tools. To address this challenge, the Sentiment Analysis on Arabic Dialects in…

Computation and Language · Computer Science 2025-11-18 Maram Alharbi , Salmane Chafik , Saad Ezzini , Ruslan Mitkov , Tharindu Ranasinghe , Hansi Hettiarachchi

ASR has achieved remarkable global progress, yet African low-resource languages remain rigorously underrepresented, producing barriers to digital inclusion across the continent with more than +2000 languages. This systematic literature…

State-of-the-art natural language processing systems rely on supervision in the form of annotated data to learn competent models. These models are generally trained on data in a single language (usually English), and cannot be directly used…

Computation and Language · Computer Science 2018-09-14 Alexis Conneau , Guillaume Lample , Ruty Rinott , Adina Williams , Samuel R. Bowman , Holger Schwenk , Veselin Stoyanov

Several widely used software applications involve some form of processing of natural language, with tasks ranging from digitising hardcopies and text processing to speech generation. Varied language resources are used to develop software…

Computation and Language · Computer Science 2026-01-21 C. Maria Keet , Langa Khumalo

The NLP pipeline has evolved dramatically in the last few years. The first step in the pipeline is to find suitable annotated datasets to evaluate the tasks we are trying to solve. Unfortunately, most of the published datasets lack metadata…

Computation and Language · Computer Science 2021-10-14 Zaid Alyafeai , Maraim Masoud , Mustafa Ghaleb , Maged S. Al-shaibani

Figures of Speech (FoS) consist of multi-word phrases that are deeply intertwined with culture. While Neural Machine Translation (NMT) performs relatively well with the figurative expressions of high-resource languages, it often faces…

Computation and Language · Computer Science 2026-02-11 Johan Sofalas , Dilushri Pavithra , Nevidu Jayatilleke , Ruvan Weerasinghe

Existing data-to-text generation datasets are mostly limited to English. To address this lack of data, we create Table-to-Text in African languages (TaTa), the first large multilingual table-to-text dataset with a focus on African…

Computation and Language · Computer Science 2022-11-02 Sebastian Gehrmann , Sebastian Ruder , Vitaly Nikolaev , Jan A. Botha , Michael Chavinda , Ankur Parikh , Clara Rivera

This paper describes a web-based corpus of global language use with a focus on how this corpus can be used for data-driven language mapping. First, the corpus provides a representation of where national varieties of major languages are used…

Computation and Language · Computer Science 2020-04-03 Jonathan Dunn

In this paper we present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger. These two datasets were made available for the IWSLT 2022 low-resource speech translation track, and they consist of collections of…

Computation and Language · Computer Science 2022-04-12 Marcely Zanon Boito , Fethi Bougares , Florentin Barbier , Souhir Gahbiche , Loïc Barrault , Mickael Rouvier , Yannick Estève

High-quality data resources play a crucial role in learning large language models (LLMs), particularly for low-resource languages like Cantonese. Despite having more than 85 million native speakers, Cantonese is still considered a…

Computation and Language · Computer Science 2025-03-06 Jiyue Jiang , Alfred Kar Yin Truong , Yanyu Chen , Qinghang Bao , Sheng Wang , Pengan Chen , Jiuming Wang , Lingpeng Kong , Yu Li , Chuan Wu
‹ Prev 1 4 5 6 7 8 10 Next ›