中文
相关论文

相关论文: Yankari: A Monolingual Yoruba Dataset

200 篇论文

Natural language processing (NLP) has made significant progress for well-resourced languages such as English but lagged behind for low-resource languages like Setswana. This paper addresses this gap by presenting PuoBERTa, a customised…

计算与语言 · 计算机科学 2023-10-25 Vukosi Marivate , Moseli Mots'Oehli , Valencia Wagner , Richard Lastrucci , Isheanesu Dzingirai

Natural Language Generation (NLG) for non-English languages is hampered by the scarcity of datasets in these languages. In this paper, we present the IndicNLG Benchmark, a collection of datasets for benchmarking NLG for 11 Indic languages.…

Comprehensive monolingual Natural Language Processing (NLP) surveys are essential for assessing language-specific challenges, resource availability, and research gaps. However, existing surveys often lack standardized methodologies, leading…

计算与语言 · 计算机科学 2025-06-19 Juli Bakagianni , Kanella Pouli , Maria Gavriilidou , John Pavlopoulos

In this work, we present BanglaParaphrase, a high-quality synthetic Bangla Paraphrase dataset curated by a novel filtering pipeline. We aim to take a step towards alleviating the low resource status of the Bangla language in the NLP domain…

计算与语言 · 计算机科学 2022-10-12 Ajwad Akil , Najrin Sultana , Abhik Bhattacharjee , Rifat Shahriyar

If today some African languages like Swahili have enough resources to develop high-performing Natural Language Processing (NLP) systems, many other languages spoken on the continent are still lacking such support. For these languages, still…

计算与语言 · 计算机科学 2024-12-19 Naira Abdou Mohamed , Zakarya Erraji , Abdessalam Bahafid , Imade Benelallam

Large Language Models (LLMs) like GPT-4 and LLaMA have shown incredible proficiency at natural language processing tasks and have even begun to excel at tasks across other modalities such as vision and audio. Despite their success, LLMs…

计算与语言 · 计算机科学 2024-03-12 Michael Andersland

Yor\`ub\'a is a widely spoken West African language with a writing system rich in tonal and orthographic diacritics. With very few exceptions, diacritics are omitted from electronic texts, due to limited device and application support.…

计算与语言 · 计算机科学 2018-10-31 Iroro Orife

Stereotype repositories are critical to assess generative AI model safety, but currently lack adequate global coverage. It is imperative to prioritize targeted expansion, strategically addressing existing deficits, over merely increasing…

We present mahaNLP, an open-source natural language processing (NLP) library specifically built for the Marathi language. It aims to enhance the support for the low-resource Indian language Marathi in the field of NLP. It is an easy-to-use,…

计算与语言 · 计算机科学 2023-11-07 Vidula Magdum , Omkar Dhekane , Sharayu Hiwarkhedkar , Saloni Mittal , Raviraj Joshi

Bengali is an underrepresented language in NLP research. However, it remains a challenge due to its unique linguistic structure and computational constraints. In this work, we systematically investigate the challenges that hinder Bengali…

计算与语言 · 计算机科学 2025-08-01 Shimanto Bhowmik , Tawsif Tashwar Dipto , Md Sazzad Islam , Sheryl Hsu , Tahsin Reasat

NLP research is impeded by a lack of resources and awareness of the challenges presented by underrepresented languages and dialects. Focusing on the languages spoken in Indonesia, the second most linguistically diverse and the fourth most…

State-of-the-art natural language processing systems rely on supervision in the form of annotated data to learn competent models. These models are generally trained on data in a single language (usually English), and cannot be directly used…

This work supports further development of language technology for the languages of Africa by providing a Wikidata-derived resource of name lists corresponding to common entity types (person, location, and organization). While we are not the…

计算与语言 · 计算机科学 2021-04-02 Jonne Sälevä , Constantine Lignos

In Indonesia, local languages play an integral role in the culture. However, the available Indonesian language resources still fall into the category of limited data in the Natural Language Processing (NLP) field. This is become problematic…

计算与语言 · 计算机科学 2024-04-02 Joanito Agili Lopo , Radius Tanone

The proliferation of online offensive language necessitates the development of effective detection mechanisms, especially in multilingual contexts. This study addresses the challenge by developing and introducing novel datasets for…

计算与语言 · 计算机科学 2024-06-07 Saminu Mohammad Aliyu , Gregory Maksha Wajiga , Muhammad Murtala

Marathi is one of the most widely used languages in the world. One might expect that the latest advances in NLP research in languages like English reach such a large community. However, NLP advancements in English didn't immediately reach…

计算与语言 · 计算机科学 2024-12-25 Asang Dani , Shailesh R Sathe

The absence of explicitly tailored, accessible annotated datasets for educational purposes presents a notable obstacle for NLP tasks in languages with limited resources.This study initially explores the feasibility of using machine…

计算与语言 · 计算机科学 2024-04-29 Hailay Teklehaymanot , Dren Fazlija , Niloy Ganguly , Gourab K. Patro , Wolfgang Nejdl

This survey provides a comprehensive catalog of publicly available text and speech resources for two West African languages: Hausa, an Afroasiatic language with approximately 80-100 million speakers, and Fongbe, a Niger-Congo language…

计算与语言 · 计算机科学 2026-05-25 Mahounan Pericles Adjovi , Victor Olufemi , Roald Eiselen , Prasenjit Mitra

Recent advances in Natural Language Processing (NLP) have underscored the crucial role of high-quality datasets in building large language models (LLMs). However, while extensive resources and analyses exist for English, the landscape for…

计算与语言 · 计算机科学 2025-10-16 Dasol Choi , Woomyoung Park , Youngsook Song