中文
相关论文

相关论文: Monolingual and Parallel Corpora for Kangri Low Re…

200 篇论文

As an Indo-Aryan language with limited available data, Chakma remains largely underrepresented in language models. In this work, we introduce a novel corpus of contextually coherent Bangla-transliterated Chakma, curated from Chakma…

计算与语言 · 计算机科学 2025-11-27 Adity Khisa , Nusrat Jahan Lia , Tasnim Mahfuz Nafis , Zarif Masud , Tanzir Pial , Shebuti Rayana , Ahmedul Kabir

Despite dramatic recent progress in NLP, it is still a major challenge to apply Large Language Models (LLM) to low-resource languages. This is made visible in benchmarks such as Cross-Lingual Natural Language Inference (XNLI), a key task…

计算与语言 · 计算机科学 2025-04-15 Aung Kyaw Htet , Mark Dras

Text normalization is a crucial technology for low-resource languages which lack rigid spelling conventions or that have undergone multiple spelling reforms. Low-resource text normalization has so far relied upon hand-crafted rules, which…

计算与语言 · 计算机科学 2023-12-25 Stefano Lusito , Edoardo Ferrante , Jean Maillard

Low-resource languages present unique challenges to (neural) machine translation. We discuss the case of Bambara, a Mande language for which training data is scarce and requires significant amounts of pre-processing. More than the…

We present a new release of the Czech-English parallel corpus CzEng 2.0 consisting of over 2 billion words (2 "gigawords") in each language. The corpus contains document-level information and is filtered with several techniques to lower the…

计算与语言 · 计算机科学 2020-07-08 Tom Kocmi , Martin Popel , Ondrej Bojar

For any deep computational processing of language we need evidences, and one such set of evidences is corpus. This paper describes the development of a text-based corpus for the Bishnupriya Manipuri language. A Corpus is considered as a…

计算与语言 · 计算机科学 2013-12-12 Nayan Jyoti Kalita , Navanath Saharia , Smriti Kumar Sinha

While Indic NLP has made rapid advances recently in terms of the availability of corpora and pre-trained models, benchmark datasets on standard NLU tasks are limited. To this end, we introduce IndicXNLI, an NLI dataset for 11 Indic…

计算与语言 · 计算机科学 2022-04-20 Divyanshu Aggarwal , Vivek Gupta , Anoop Kunchukuttan

Bangla -- ranked as the 6th most widely spoken language across the world (https://www.ethnologue.com/guides/ethnologue200), with 230 million native speakers -- is still considered as a low-resource language in the natural language…

计算与语言 · 计算机科学 2021-07-27 Firoj Alam , Arid Hasan , Tanvirul Alam , Akib Khan , Janntatul Tajrin , Naira Khan , Shammur Absar Chowdhury

This paper evaluates the performance of Large Multimodal Models (LMMs) on Optical Character Recognition (OCR) in the low-resource Pashto language. Natural Language Processing (NLP) in Pashto faces several challenges due to the cursive…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Ijazul Haq , Yingjie Zhang , Irfan Ali Khan

Indian language machine translation performance is hampered due to the lack of large scale multi-lingual sentence aligned corpora and robust benchmarks. Through this paper, we provide and analyse an automated framework to obtain such a…

计算与语言 · 计算机科学 2020-11-05 Jerin Philip , Shashank Siripragada , Vinay P. Namboodiri , C. V. Jawahar

This work contributes towards balancing the inclusivity and global applicability of natural language processing techniques by proposing the first 'name entity recognition' dataset for Kurdish Sorani, a low-resource and under-represented…

计算与语言 · 计算机科学 2025-12-01 Bakhtawar Abdalla , Rebwar Mala Nabi , Hassan Eshkiki , Fabio Caraffini

In this paper we describe our efforts to make a bidirectional Congolese Swahili (SWC) to French (FRA) neural machine translation system with the motivation of improving humanitarian translation workflows. For training, we created a…

计算与语言 · 计算机科学 2021-03-22 Alp Öktem , Eric DeLuca , Rodrigue Bashizi , Eric Paquin , Grace Tang

MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual dataset we have built for the WSDM 2023 Cup challenge that focuses on ad hoc retrieval across 18 different languages, which collectively encompass…

We present a mobile and desktop keyboard suite for Idu Mishmi, an endangered Trans-Himalayan language spoken by approximately 11,000 people in Arunachal Pradesh, India. Although a Latin-based orthography was developed in 2018, no digital…

计算与语言 · 计算机科学 2026-02-24 Akhilesh Kakolu Ramarao

Large Language Models (LLMs) consistently under perform in low-resource linguistic contexts such as Konkani. This performance deficit stems from acute training data scarcity compounded by high script diversity across Devanagari, Romi and…

计算与语言 · 计算机科学 2026-03-26 Reuben Chagas Fernandes , Gaurang S. Patkar

Despite rapid advances in large language models (LLMs), low-resource languages remain excluded from NLP, limiting digital access for millions. We present PunGPT2, the first fully open-source Punjabi generative model suite, trained on a 35GB…

计算与语言 · 计算机科学 2025-10-06 Jaskaranjeet Singh , Rakesh Thakur

Aligned audio corpora are fundamental to NLP technologies such as ASR and speech translation, yet they remain scarce for underrepresented languages, hindering their technological integration. This paper introduces a methodology for…

计算与语言 · 计算机科学 2026-03-11 Samy Ouzerrout

Code-mixing is the phenomenon of using more than one language in a sentence. It is a very frequently observed pattern of communication on social media platforms. Flexibility to use multiple languages in one text message might help to…

计算与语言 · 计算机科学 2020-04-21 Vivek Srivastava , Mayank Singh

Despite much progress in recent years, the vast majority of work in natural language processing (NLP) is on standard languages with many speakers. In this work, we instead focus on low-resource languages and in particular non-standardized…

计算与语言 · 计算机科学 2023-04-20 Verena Blaschke , Hinrich Schütze , Barbara Plank

We present the Thiomi Dataset, a large-scale multimodal corpus spanning ten African languages across four language families: Swahili, Kikuyu, Kamba, Kimeru, Luo, Maasai, Kipsigis, Somali (East Africa); Wolof (West Africa); and Fulani…

计算与语言 · 计算机科学 2026-04-07 Hillary Mutisya , John Mugane , Gavin Nyamboga , Brian Chege , Maryruth Gathoni