English
Related papers

Related papers: The first large scale collection of diverse Hausa …

200 papers

As machine learning and data science applications grow ever more prevalent, there is an increased focus on data sharing and open data initiatives, particularly in the context of the African continent. Many argue that data sharing can…

Computers and Society · Computer Science 2021-03-02 Rediet Abebe , Kehinde Aruleba , Abeba Birhane , Sara Kingsley , George Obaido , Sekou L. Remy , Swathi Sadagopan

The Fon language, spoken by an average 2 million of people, is a truly low-resourced African language, with a limited online presence, and existing datasets (just to name but a few). Multitask learning is a learning paradigm that aims to…

Computation and Language · Computer Science 2023-09-13 Bonaventure F. P. Dossou , Iffanice Houndayi , Pamely Zantou , Gilles Hacheme

Natural language processing, as a data analytics related technology, is used widely in many research areas such as artificial intelligence, human language processing, and translation. At present, due to explosive growth of data, there are…

Computation and Language · Computer Science 2016-08-17 Emre Erturk , Hong Shi

Language models are the foundation of current neural network-based models for natural language understanding and generation. However, research on the intrinsic performance of language models on African languages has been extremely limited,…

Computation and Language · Computer Science 2021-04-05 Stuart Mesham , Luc Hayward , Jared Shapiro , Jan Buys

This paper introduces a centralized, open-source dataset repository designed to advance NLP and NMT for Assamese, a low-resource language. The repository, available at GitHub, supports various tasks like sentiment analysis, named entity…

Computation and Language · Computer Science 2024-10-17 S. Tamang , D. J. Bora

While resources for English language are fairly sufficient to understand content on social media, similar resources in Arabic are still immature. The main reason that the resources in Arabic are insufficient is that Arabic has many dialects…

Computation and Language · Computer Science 2023-09-22 Fatimah Alzamzami , Abdulmotaleb El Saddik

We present an analysis pipeline and best practice guidelines for building and curating corpora of everyday conversation in diverse languages. Surveying language documentation corpora and other resources that cover 67 languages and varieties…

Computation and Language · Computer Science 2022-05-11 Andreas Liesenfeld , Mark Dingemanse

Multilingual NLP often relies on dataset counts from centralized catalogues to characterize which languages are resource-rich or resource-poor. However, these catalogues record only one layer of dataset visibility: what has been registered…

Computation and Language · Computer Science 2026-05-19 Zhiyin Tan , Changxu Duan

Australian Aboriginal languages are of significant cultural and linguistic value but remain severely underrepresented in modern speech AI systems. While state-of-the-art speech foundation models and automatic speech recognition excel in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Ting Dang , Trini Manoj Jeyaseelan , Eliathamby Ambikairajah , Vidhyasaharan Sethu

This paper introduces AfriHG -- a news headline generation dataset created by combining from XLSum and MasakhaNEWS datasets focusing on 16 languages widely spoken by Africa. We experimented with two seq2eq models (mT5-base and AfriTeVa V2),…

Computation and Language · Computer Science 2024-12-31 Toyib Ogunremi , Serah Akojenu , Anthony Soronnadi , Olubayo Adekanmbi , David Ifeoluwa Adelani

High-resource language models often fall short in the African context, where there is a critical need for models that are efficient, accessible, and locally relevant, even amidst significant computing and data constraints. This paper…

Currently, natural language processing (NLP) models proliferate language discrimination leading to potentially harmful societal impacts as a result of biased outcomes. For example, part-of-speech taggers trained on Mainstream American…

Computation and Language · Computer Science 2022-06-22 Jamell Dacon

Linguistic disparity in the NLP world is a problem that has been widely acknowledged recently. However, different facets of this problem, or the reasons behind this disparity are seldom discussed within the NLP community. This paper…

Computation and Language · Computer Science 2022-10-21 Surangika Ranathunga , Nisansa de Silva

Large Language Models (LLMs) have rapidly increased in size and apparent capabilities in the last three years, but their training data is largely English text. There is growing interest in multilingual LLMs, and various efforts are striving…

Social media platforms have become central to global communication, yet they also facilitate the spread of hate speech. For underrepresented dialects like Levantine Arabic, detecting hate speech presents unique cultural, ethical, and…

Computation and Language · Computer Science 2024-12-17 Ahmed Haj Ahmed , Rui-Jie Yew , Xerxes Minocher , Suresh Venkatasubramanian

Creating spoken dialogue datasets is methodologically challenging, and these challenges are amplified when the goal is to build multilingual, multi-parallel datasets at scale. This work introduces HEALTHDIAL, a large-scale, multilingual,…

Computation and Language · Computer Science 2026-05-29 Songbo Hu , Yinhong Liu , Ej Zhou , Evgeniia Razumovskaia , Xiaobin Wang , Alexander Fraser , Ivan Vulić , Anna Korhonen

Human language is firstly spoken and only secondarily written. Text, however, is a very convenient and efficient representation of language, and modern civilization has made it ubiquitous. Thus the field of NLP has overwhelmingly focused on…

Computation and Language · Computer Science 2023-05-24 Grzegorz Chrupała

The Serbian language is a Slavic language spoken by over 12 million speakers and well understood by over 15 million people. In the area of natural language processing, it can be considered a low-resourced language. Also, Serbian is…

Computation and Language · Computer Science 2023-04-13 Ulfeta A. Marovac , Aldina R. Avdić , Nikola Lj. Milošević

As language and speech technologies become more advanced, the lack of fundamental digital resources for African languages, such as data, spell checkers and Part of Speech taggers, means that the digital divide between these languages and…

Computation and Language · Computer Science 2020-07-24 Kathleen Siminyu , Sackey Freshia , Jade Abbott , Vukosi Marivate
‹ Prev 1 8 9 10 Next ›