English
Related papers

Related papers: Language Diversity: Evaluating Language Usage and …

200 papers

Multimodal AI research has overwhelmingly focused on high-resource languages, hindering the democratization of advancements in the field. To address this, we present AfriCaption, a comprehensive framework for multilingual image captioning…

Computation and Language · Computer Science 2025-10-21 Mardiyyah Oduwole , Prince Mireku , Fatimo Adebanjo , Oluwatosin Olajide , Mahi Aminu Aliyu , Jekaterina Novikova

We present the findings of SemEval-2023 Task 12, a shared task on sentiment analysis for low-resource African languages using Twitter dataset. The task featured three subtasks; subtask A is monolingual sentiment classification with 12…

Large Language Models (LLMs) have demonstrated remarkable capabilities in reasoning tasks, leading to their widespread deployment. However, recent studies have highlighted concerning biases in these models, particularly in their handling of…

Computation and Language · Computer Science 2025-03-07 Runtao Zhou , Guangya Wan , Saadia Gabriel , Sheng Li , Alexander J Gates , Maarten Sap , Thomas Hartvigsen

It is a well-known fact that current AI-based language technology -- language models, machine translation systems, multilingual dictionaries and corpora -- focuses on the world's 2-3% most widely spoken languages. Recent research efforts…

Computation and Language · Computer Science 2023-07-26 Gábor Bella , Paula Helm , Gertraud Koch , Fausto Giunchiglia

Existing AI bias evaluation benchmarks largely reflect Western perspectives, leaving African contexts underrepresented and enabling harmful stereotypes in applications across various domains. To address this gap, we introduce AfriStereo,…

Computation and Language · Computer Science 2025-12-01 Yann Le Beux , Oluchi Audu , Oche D. Ankeli , Dhananjay Balakrishnan , Melissah Weya , Marie D. Ralaiarinosy , Ignatius Ezeani

Progress in AI is driven largely by the scale and quality of training data. Despite this, there is a deficit of empirical analysis examining the attributes of well-established datasets beyond text. In this work we conduct the largest and…

Nigeria is a multilingual country with 500+ languages. Naija is a Nigerian Pidgin spoken by approximately 120M speakers and it is a mixed language (e.g., English, Portuguese, Yoruba, Hausa and Igbo). Although it has mainly been a spoken…

Computation and Language · Computer Science 2025-05-01 David Ifeoluwa Adelani , A. Seza Doğruöz , Iyanuoluwa Shode , Anuoluwapo Aremu

While many speakers of low-resource languages regularly code-switch between their languages and other regional languages or English, datasets of codeswitched speech are too small to train bespoke acoustic models from scratch or do language…

Computation and Language · Computer Science 2023-11-28 Tolúlopé Ògúnrèmí , Christopher D. Manning , Dan Jurafsky

Multilingual pretrained language models (mPLMs) acquire valuable, generalizable linguistic information during pretraining and have advanced the state of the art on task-specific finetuning. To date, only ~31 out of ~2,000 African languages…

Computation and Language · Computer Science 2023-05-30 Ife Adebara , AbdelRahim Elmadany , Muhammad Abdul-Mageed , Alcides Alcoba Inciarte

Crafting an effective Automatic Speech Recognition (ASR) solution for dialects demands innovative approaches that not only address the data scarcity issue but also navigate the intricacies of linguistic diversity. In this paper, we address…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-26 Ahmed Amine Ben Abdallah , Ata Kabboudi , Amir Kanoun , Salah Zaiem

The Internet is currently the largest platform for global communication including expressions of opinions, reviews, contents, images, videos and so forth. Moreover, social media has now become a very broad and highly engaging platform due…

Computation and Language · Computer Science 2024-01-17 Sristy Shidul Nath , Razuan Karim , Mahdi H. Miraz

Indigenous African languages are categorized as under-served in Natural Language Processing. They therefore experience poor digital inclusivity and information access. The processing challenge with such languages has been how to use machine…

Computation and Language · Computer Science 2025-01-17 Barack Wanjawa , Lilian Wanzare , Florence Indede , Owen McOnyango , Edward Ombui , Lawrence Muchemi

Despite comprising one-third of global languages, African languages are critically underrepresented in Artificial Intelligence (AI), threatening linguistic diversity and cultural heritage. Ghanaian languages, in particular, face an alarming…

Computation and Language · Computer Science 2024-05-14 Sheriff Issaka , Zhaoyi Zhang , Mihir Heda , Keyi Wang , Yinka Ajibola , Ryan DeMar , Xuefeng Du

In this work, we present AfriHuBERT, an extension of mHuBERT-147, a compact self-supervised learning (SSL) model pretrained on 147 languages. While mHuBERT-147 covered 16 African languages, we expand this to 1,226 through continued…

Computation and Language · Computer Science 2025-06-03 Jesujoba O. Alabi , Xuechen Liu , Dietrich Klakow , Junichi Yamagishi

The rapid expansion of social media platforms has significantly increased the dissemination of forged content and misinformation, making the detection of fake news a critical area of research. Although fact-checking efforts predominantly…

Computation and Language · Computer Science 2025-06-03 Muhammad Islam , Javed Ali Khan , Mohammed Abaker , Ali Daud , Azeem Irshad

Developing robust automatic speech recognition (ASR) systems for Arabic requires effective strategies to manage its diversity. Existing ASR systems mainly cover the modern standard Arabic (MSA) variety and few high-resource dialects, but…

Computation and Language · Computer Science 2025-06-02 Amirbek Djanibekov , Hawau Olamide Toyin , Raghad Alshalan , Abdullah Alitr , Hanan Aldarmaki

The training data for LLMs embeds societal values, increasing their familiarity with the language's culture. Our analysis found that 44% of the variance in the ability of GPT-4o to reflect the societal values of a country, as measured by…

Computation and Language · Computer Science 2024-10-15 Sharif Kazemi , Gloria Gerhardt , Jonty Katz , Caroline Ida Kuria , Estelle Pan , Umang Prabhakar

Language models built from various sources are the foundation of today's NLP progress. However, for many low-resource languages, the diversity of domains is often limited, more biased to a religious domain, which impacts their performance…

The hidden nature and the limited accessibility of the Dark Web, combined with the lack of public datasets in this domain, make it difficult to study its inherent characteristics such as linguistic properties. Previous works on text…

Computation and Language · Computer Science 2022-05-05 Youngjin Jin , Eugene Jang , Yongjae Lee , Seungwon Shin , Jin-Woo Chung

This research addresses the challenge of developing speech applications for zero-resource languages that lack labelled data. It specifically uses acoustic word embedding (AWE) -- fixed-dimensional representations of variable-duration speech…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-24 Christiaan Jacobs