English
Related papers

Related papers: AfroDigits: A Community-Driven Spoken Digit Datase…

200 papers

Sentiment analysis is a fundamental and valuable task in NLP. However, due to limitations in data and technological availability, research into sentiment analysis of African languages has been fragmented and lacking. With the recent release…

Computation and Language · Computer Science 2023-10-24 Saurav K. Aryal , Howard Prioleau , Surakshya Aryal

Idioms are figurative expressions whose meanings often cannot be inferred from their individual words, making them difficult to process computationally and posing challenges for human experimental studies. This survey reviews datasets…

Computation and Language · Computer Science 2025-08-19 Michael Flor , Xinyi Liu , Anna Feldman

Personalization of speech models on mobile devices (on-device personalization) is an active area of research, but more often than not, mobile devices have more text-only data than paired audio-text data. We explore training a personalized…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-05 Theresa Breiner , Swaroop Ramaswamy , Ehsan Variani , Shefali Garg , Rajiv Mathews , Khe Chai Sim , Kilol Gupta , Mingqing Chen , Lara McConnaughey

Developing Automatic Speech Recognition (ASR) systems for Tunisian Arabic Dialect is challenging due to the dialect's linguistic complexity and the scarcity of annotated speech datasets. To address these challenges, we propose the LinTO…

Computation and Language · Computer Science 2025-04-04 Hedi Naouara , Jean-Pierre Lorré , Jérôme Louradour

We present the findings of SemEval-2023 Task 12, a shared task on sentiment analysis for low-resource African languages using Twitter dataset. The task featured three subtasks; subtask A is monolingual sentiment classification with 12…

We present the Multilingual Cloud Corpus, the first national-scale, parallel, multimodal linguistic dataset of Bangladesh's ethnic and indigenous languages. Despite being home to approximately 40 minority languages spanning four language…

Computation and Language · Computer Science 2026-03-09 Mohammad Mamun Or Rashid

Despite advancements in conversational AI, language models encounter challenges to handle diverse conversational tasks, and existing dialogue dataset collections often lack diversity and comprehensiveness. To tackle these issues, we…

Computation and Language · Computer Science 2024-02-06 Jianguo Zhang , Kun Qian , Zhiwei Liu , Shelby Heinecke , Rui Meng , Ye Liu , Zhou Yu , Huan Wang , Silvio Savarese , Caiming Xiong

Assessing the veracity of a claim made online is a complex and important task with real-world implications. When these claims are directed at communities with limited access to information and the content concerns issues such as healthcare…

In this work, we present AfriNLLB, a series of lightweight models for efficient translation from and into African languages. AfriNLLB supports 15 language pairs (30 translation directions), including Swahili, Hausa, Yoruba, Amharic, Somali,…

Computation and Language · Computer Science 2026-02-11 Yasmin Moslem , Aman Kassahun Wassie , Amanuel Gizachew Abebe

To benchmark Bengali digit recognition algorithms, a large publicly available dataset is required which is free from biases originating from geographical location, gender, and age. With this aim in mind, NumtaDB, a dataset consisting of…

Computer Vision and Pattern Recognition · Computer Science 2018-06-08 Samiul Alam , Tahsin Reasat , Rashed Mohammad Doha , Ahmed Imtiaz Humayun

Benchmarking plays a pivotal role in assessing and enhancing the performance of compact deep learning models designed for execution on resource-constrained devices, such as microcontrollers. Our study introduces a novel, entirely…

Sound · Computer Science 2024-03-18 René Groh , Nina Goes , Andreas M. Kist

Real-time speech assistants are becoming increasingly popular for ensuring improved accessibility to information. Bengali, being a low-resource language with a high regional dialectal diversity, has seen limited progress in developing such…

Computation and Language · Computer Science 2025-11-17 Jakir Hasan , Shubhashis Roy Dipta

Text embeddings are an essential building component of several NLP tasks such as retrieval-augmented generation which is crucial for preventing hallucinations in LLMs. Despite the recent release of massively multilingual MTEB (MMTEB),…

Computation and Language · Computer Science 2026-03-09 Kosei Uemura , Miaoran Zhang , David Ifeoluwa Adelani

Ethiopic/Amharic script is one of the oldest African writing systems, which serves at least 23 languages (e.g., Amharic, Tigrinya) in East Africa for more than 120 million people. The Amharic writing system, Abugida, has 282 syllables, 15…

Computer Vision and Pattern Recognition · Computer Science 2022-04-12 Wondimu Dikubab , Dingkang Liang , Minghui Liao , Xiang Bai

Useful conversational agents must accurately capture named entities to minimize error for downstream tasks, for example, asking a voice assistant to play a track from a certain artist, initiating navigation to a specific location, or…

Language is a method by which individuals express their thoughts. Each language has its own set of alphabetic and numeric characters. People can communicate with one another through either oral or written communication. However, each…

One of the major challenges for developing automatic speech recognition (ASR) for low-resource languages is the limited access to labeled data with domain-specific variations. In this study, we propose a pseudo-labeling approach to develop…

Slot-filling and intent detection are well-established tasks in Conversational AI. However, current large-scale benchmarks for these tasks often exclude evaluations of low-resource languages and rely on translations from English benchmarks,…

Large language models (LLMs) are increasingly multilingual, yet open models continue to underperform relative to proprietary systems, with the gap most pronounced for African languages. Continued pre-training (CPT) offers a practical route…

Computation and Language · Computer Science 2026-05-06 Hao Yu , Tianyi Xu , Michael A. Hedderich , Wassim Hamidouche , Syed Waqas Zamir , David Ifeoluwa Adelani

The primary obstacle to developing technologies for low-resource languages is the lack of usable data. In this paper, we report the adoption and deployment of 4 technology-driven methods of data collection for Gondi, a low-resource…