English
Related papers

Related papers: NusaCrowd: Open Source Initiative for Indonesian N…

200 papers

Weak supervision has emerged as a promising approach for rapid and large-scale dataset creation in response to the increasing demand for accelerated NLP development. By leveraging labeling functions, weak supervision allows practitioners to…

Computation and Language · Computer Science 2023-10-25 Mega Fransiska , Diah Pitaloka , Saripudin , Satrio Putra , Lintang Sutawika

Figurative language permeates human communication, but at the same time is relatively understudied in NLP. Datasets have been created in English to accelerate progress towards measuring and improving figurative language processing in…

Hausa Natural Language Processing (NLP) has gained increasing attention in recent years, yet remains understudied as a low-resource language despite having over 120 million first-language (L1) and 80 million second-language (L2) speakers…

Culturally grounded commonsense reasoning is underexplored in low-resource languages due to scarce data and costly native annotation. We test whether large language models (LLMs) can generate culturally nuanced narratives for such settings.…

Computation and Language · Computer Science 2025-09-12 Salsabila Zahirah Pranida , Rifo Ahmad Genadi , Fajri Koto

Natural Language Processing (NLP) for low-resource languages remains fundamentally constrained by the lack of textual corpora, standardized orthographies, and scalable annotation pipelines. While recent advances in large language models…

Computation and Language · Computer Science 2026-02-10 Bonaventure F. P. Dossou , Henri Aïdasso

In its daily use, the Indonesian language is riddled with informality, that is, deviations from the standard in terms of vocabulary, spelling, and word order. On the other hand, current available Indonesian NLP models are typically…

Indonesian is an agglutinative language since it has a compounding process of word-formation. Therefore, the translation model of this language requires a mechanism that is even lower than the word level, referred to as the sub-word level.…

Computation and Language · Computer Science 2022-07-04 Mukhlis Amien , Feng Chong , Huang Heyan

In this paper, we introduce a large-scale Indonesian summarization dataset. We harvest articles from Liputan6.com, an online news portal, and obtain 215,827 document-summary pairs. We leverage pre-trained language models to develop…

Computation and Language · Computer Science 2020-11-03 Fajri Koto , Jey Han Lau , Timothy Baldwin

Although region-specific large language models (LLMs) are increasingly developed, their safety remains underexplored, particularly in culturally diverse settings like Indonesia, where sensitivity to local norms is essential and highly…

Computation and Language · Computer Science 2025-06-04 Muhammad Falensi Azmi , Muhammad Dehan Al Kautsar , Alfan Farizki Wicaksono , Fajri Koto

Large Language Models (LLMs) are increasingly being used to generate synthetic data for training and evaluating models. However, it is unclear whether they can generate a good quality of question answering (QA) dataset that incorporates…

Computation and Language · Computer Science 2024-10-08 Rifki Afina Putri , Faiz Ghifari Haznitrama , Dea Adhista , Alice Oh

This paper introduces a centralized, open-source dataset repository designed to advance NLP and NMT for Assamese, a low-resource language. The repository, available at GitHub, supports various tasks like sentiment analysis, named entity…

Computation and Language · Computer Science 2024-10-17 S. Tamang , D. J. Bora

This paper proposes the creation of a Swahili Question Answering (QA) benchmark dataset, aimed at addressing the underrepresentation of Swahili in natural language processing (NLP). Drawing from established benchmarks like SQuAD, GLUE,…

Computation and Language · Computer Science 2024-10-21 Alfred Malengo Kondoro

Performance of NLP systems is typically evaluated by collecting a large-scale dataset by means of crowd-sourcing to train a data-driven model and evaluate it on a held-out portion of the data. This approach has been shown to suffer from…

Computation and Language · Computer Science 2024-08-12 Viktor Schlegel , Goran Nenadic , Riza Batista-Navarro

Massively multilingual neural machine translation (MMNMT) has been proven to enhance the translation quality of low-resource languages. In this paper, we empirically investigate the translation robustness of Indonesian-Chinese translation…

Computation and Language · Computer Science 2024-05-14 Supryadi , Leiyu Pan , Deyi Xiong

Crowdsourcing is widely used to create data for common natural language understanding tasks. Despite the importance of these datasets for measuring and refining model understanding of language, there has been little focus on the…

Computation and Language · Computer Science 2021-06-03 Nikita Nangia , Saku Sugawara , Harsh Trivedi , Alex Warstadt , Clara Vania , Samuel R. Bowman

Despite the progress we have recorded in the last few years in multilingual natural language processing, evaluation is typically limited to a small set of languages with available datasets which excludes a large number of low-resource…

Computation and Language · Computer Science 2024-03-08 David Ifeoluwa Adelani , Hannah Liu , Xiaoyu Shen , Nikita Vassilyev , Jesujoba O. Alabi , Yanke Mao , Haonan Gao , Annie En-Shiun Lee

The transparency nature of Open Data is beneficial for citizens to evaluate government work performance. In Indonesia, each government bodies or ministry have their own standard operating procedure on data treatment resulting in incoherent…

Computers and Society · Computer Science 2021-02-18 A. Alamsyah , T. T. Gustyana , A. D. Fajaryanto , D. Septiafani

Sentence simplification aims to make complex text more accessible by reducing linguistic complexity while preserving the original meaning. However, progress in this area remains limited for mid-resource and low-resource languages due to the…

Faced with a considerable lack of resources in African languages to carry out work in Natural Language Processing (NLP), Natural Language Understanding (NLU) and artificial intelligence, the research teams of NTeALan association has set…

Computation and Language · Computer Science 2021-04-01 Elvis Mboning Tchiaze