中文
相关论文

相关论文: Correcting FLORES Evaluation Dataset for Four Afri…

200 篇论文

We present the first self-supervised multilingual speech model trained exclusively on African speech. The model learned from nearly 60 000 hours of unlabeled speech segments in 21 languages and dialects spoken in sub-Saharan Africa. On the…

计算与语言 · 计算机科学 2024-04-23 Antoine Caubrière , Elodie Gauthier

AfriSenti-SemEval Shared Task 12 of SemEval-2023. The task aims to perform monolingual sentiment classification (sub-task A) for 12 African languages, multilingual sentiment classification (sub-task B), and zero-shot sentiment…

Neural retrieval and GPT-style generative models rely on large, high-quality supervised data, which is still scarce for low-resource languages such as Amharic. We release an Amharic data resource consisting of two datasets that supports…

计算与语言 · 计算机科学 2026-02-11 Tilahun Yeshambel , Moncef Garouani , Josiane Mothe

We address a notable gap in Natural Language Processing (NLP) by introducing a collection of resources designed to improve Machine Translation (MT) for low-resource languages, with a specific focus on African languages. First, we introduce…

计算与语言 · 计算机科学 2024-07-15 AbdelRahim Elmadany , Ife Adebara , Muhammad Abdul-Mageed

Large language models (LLMs) have achieved impressive results in high-resource languages like English, yet their effectiveness in low-resource and morphologically rich languages remains underexplored. In this paper, we present a…

计算与语言 · 计算机科学 2026-02-13 Chengxuan Xia , Qianye Wu , Hongbin Guan , Sixuan Tian , Yilun Hao , Xiaoyu Wu

We present the first parallel dataset for English-Tulu translation. Tulu, classified within the South Dravidian linguistic family branch, is predominantly spoken by approximately 2.5 million individuals in southwestern India. Our dataset is…

计算与语言 · 计算机科学 2024-03-29 Manu Narayanan , Noëmi Aepli

Although researchers and practitioners are pushing the boundaries and enhancing the capacities of NLP tools and methods, works on African languages are lagging. A lot of focus on well resourced languages such as English, Japanese, German,…

计算与语言 · 计算机科学 2020-04-03 Ignatius Ezeani , Paul Rayson , Ikechukwu Onyenwe , Chinedu Uchechukwu , Mark Hepple

Although LLMs have attained significant success in high-resource languages, their capacity in low-resource linguistic environments like Kannada and Arabic is not yet fully understood. This work benchmarking the performance of multilingual…

计算与语言 · 计算机科学 2025-07-29 Maitha Alshehhi , Ahmed Sharshar , Mohsen Guizani

The evolution of the Internet has increased the amount of information that is expressed by people on different platforms. This information can be product reviews, discussions on forums, or social media platforms. Accessibility of these…

计算与语言 · 计算机科学 2021-04-20 Gati L. Martin , Medard E. Mswahili , Young-Seob Jeong

Words embedding (distributed word vector representations) have become an essential component of many natural language processing (NLP) tasks such as machine translation, sentiment analysis, word analogy, named entity recognition and word…

计算与语言 · 计算机科学 2020-01-08 Idris Abdulmumin , Bashir Shehu Galadanci

We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark. FLEURS is an n-way parallel speech dataset in 102 languages built on top of the machine translation FLoRes-101 benchmark, with…

计算与语言 · 计算机科学 2022-05-26 Alexis Conneau , Min Ma , Simran Khanuja , Yu Zhang , Vera Axelrod , Siddharth Dalmia , Jason Riesa , Clara Rivera , Ankur Bapna

Given that South African education is in crisis, strategies for improvement and sustainability of high-quality, up-to-date education must be explored. In the migration of education online, inclusion of machine translation for low-resourced…

计算与语言 · 计算机科学 2018-11-15 Jade Z. Abbott , Laura Martinus

Transfer learning has led to large gains in performance for nearly all NLP tasks while making downstream models easier and faster to train. This has also been extended to low-resourced languages, with some success. We investigate the…

计算与语言 · 计算机科学 2023-09-12 Michael Beukman , Manuel Fokam

This paper presents our findings of the Multilingual Shared Task on Hallucinations and Related Observable Overgeneration Mistakes, MU-SHROOM, which focuses on identifying hallucinations and related overgeneration errors in large language…

With the advent of Deep Learning based Artificial Neural Networks models, Natural Language Processing (NLP) has witnessed significant improvements in textual data processing in terms of its efficiency and accuracy. However, the research is…

计算与语言 · 计算机科学 2023-10-05 Mubashir Munaf , Hammad Afzal , Naima Iltaf , Khawir Mahmood

Large language models (LLMs) have shown remarkable progress in reasoning abilities and general natural language processing (NLP) tasks, yet their performance on Arabic data, characterized by rich morphology, diverse dialects, and complex…

计算与语言 · 计算机科学 2025-12-16 Ahmed Hasanaath , Aisha Alansari , Ahmed Ashraf , Chafik Salmane , Hamzah Luqman , Saad Ezzini

Processing low-resource languages, such as Kiswahili, using machine learning is difficult due to lack of adequate training data. However, such low-resource languages are still important for human communication and are already in daily use…

计算与语言 · 计算机科学 2025-01-17 Barack Wamkaya Wanjawa , Lawrence Muchemi , Evans Miriti

This paper reports on the semi-supervised development of acoustic and language models for under-resourced, code-switched speech in five South African languages. Two approaches are considered. The first constructs four separate bilingual…

音频与语音处理 · 电气工程与系统科学 2020-03-09 Astik Biswas , Emre Yılmaz , Febe de Wet , Ewald van der Westhuizen , Thomas Niesler

This paper relates work done during the DiLAF project. It consists in converting 5 bilingual African language-French dictionaries originally in Word format into XML following the LMF model. The languages processed are Bambara, Hausa,…

计算与语言 · 计算机科学 2014-05-26 Chantal Enguehard , Mathieu Mangeot

This research article examines the effectiveness of various pretraining strategies for developing machine translation models tailored to low-resource languages. Although this work considers several low-resource languages, including…

计算与语言 · 计算机科学 2025-10-30 Idriss Nguepi Nguefack , Mara Finkelstein , Toadoum Sari Sakayo