English
Related papers

Related papers: NusaCrowd: Open Source Initiative for Indonesian N…

200 papers

We present IndoNLI, the first human-elicited NLI dataset for Indonesian. We adapt the data collection protocol for MNLI and collect nearly 18K sentence pairs annotated by crowd workers and experts. The expert-annotated data is used…

Computation and Language · Computer Science 2022-03-30 Rahmad Mahendra , Alham Fikri Aji , Samuel Louvan , Fahrurrozi Rahman , Clara Vania

India's linguistic landscape, spanning 22 scheduled languages and hundreds of marginalized dialects, has driven rapid growth in NLP datasets, benchmarks, and pretrained models. However, no dedicated survey consolidates resources developed…

Computation and Language · Computer Science 2026-04-21 Raghvendra Kumar , Devankar Raj , Sriparna Saha

Making use of off-the-shelf resources of resource-rich languages to transfer knowledge for low-resource languages raises much attention recently. The requirements of enabling the model to reach the reliable performance lack well guided,…

Computation and Language · Computer Science 2024-10-25 Donglin Di , Weinan Zhang , Yue Zhang , Fanglin Wang

The vast majority of the world's languages, particularly creoles like Nagamese, remain severely under-resourced in Natural Language Processing (NLP), creating a significant barrier to their representation in digital technology. This paper…

Computation and Language · Computer Science 2025-12-16 Agniva Maiti , Manya Pandey , Murari Mandal

This paper presents an overview of a program designed to address the growing need for developing freely available speech resources for under-represented languages. At present we have released 38 datasets for building text-to-speech and…

This study provides an overview of the history of the development of Natural Language Processing (NLP) in the context of the Indonesian language, with a focus on the basic technologies, methods, and practical applications that have been…

Computation and Language · Computer Science 2023-04-07 Mukhlis Amien

An ideal speech recognition model has the capability to transcribe speech accurately under various characteristics of speech signals, such as speaking style (read and spontaneous), speech context (formal and informal), and background noise…

Computation and Language · Computer Science 2024-10-15 Aulia Adila , Dessi Lestari , Ayu Purwarianti , Dipta Tanaya , Kurniawati Azizah , Sakriani Sakti

Natural Language Generation (NLG) for non-English languages is hampered by the scarcity of datasets in these languages. In this paper, we present the IndicNLG Benchmark, a collection of datasets for benchmarking NLG for 11 Indic languages.…

Computation and Language · Computer Science 2022-10-28 Aman Kumar , Himani Shrotriya , Prachi Sahu , Raj Dabre , Ratish Puduppully , Anoop Kunchukuttan , Amogh Mishra , Mitesh M. Khapra , Pratyush Kumar

We present COPAL-ID, a novel, public Indonesian language common sense reasoning dataset. Unlike the previous Indonesian COPA dataset (XCOPA-ID), COPAL-ID incorporates Indonesian local and cultural nuances, and therefore, provides a more…

Computation and Language · Computer Science 2024-04-23 Haryo Akbarianto Wibowo , Erland Hilman Fuadi , Made Nindyatama Nityasya , Radityo Eko Prasojo , Alham Fikri Aji

Machine Reading Comprehension (MRC) has become one of the essential tasks in Natural Language Understanding (NLU) as it is often included in several NLU benchmarks (Liang et al., 2020; Wilie et al., 2020). However, most MRC datasets only…

Computation and Language · Computer Science 2022-10-26 Rifki Afina Putri , Alice Oh

Multimodal learning on video and text has seen significant progress, particularly in tasks like text-to-video retrieval, video-to-text retrieval, and video captioning. However, most existing methods and datasets focus exclusively on…

Multimedia · Computer Science 2025-07-15 Willy Fitra Hendria

Large Language Models (LLMs) have demonstrated exceptional promise in translation tasks for high-resource languages. However, their performance in low-resource languages is limited by the scarcity of both parallel and monolingual corpora,…

Computation and Language · Computer Science 2024-10-11 William Tan , Kevin Zhu

Recently there have been intensifying efforts to improve the understanding of Indonesian cultures by large language models (LLMs). An attractive source of cultural knowledge that has been largely overlooked is local journals of social…

Computation and Language · Computer Science 2026-01-21 Adimulya Kartiyasa , Bao Gia Cao , Boyang Li

Hate speech poses a significant threat to social harmony. Over the past two years, Indonesia has seen a ten-fold increase in the online hate speech ratio, underscoring the urgent need for effective detection mechanisms. However, progress is…

Computation and Language · Computer Science 2025-06-13 Lucky Susanto , Musa Izzanardi Wijanarko , Prasetia Anugrah Pratama , Traci Hong , Ika Idris , Alham Fikri Aji , Derry Wijaya

Automatic speech recognition systems have achieved remarkable performance on fluent speech but continue to degrade significantly when processing stuttered speech, a limitation that is particularly acute for low-resource languages like…

Computation and Language · Computer Science 2026-01-15 Fadhil Muhammad , Alwin Djuliansah , Adrian Aryaputra Hamzah , Kurniawati Azizah

Building Natural Language Understanding (NLU) capabilities for Indic languages, which have a collective speaker base of more than one billion speakers is absolutely crucial. In this work, we aim to improve the NLU capabilities of Indic…

Computation and Language · Computer Science 2023-05-25 Sumanth Doddapaneni , Rahul Aralikatte , Gowtham Ramesh , Shreya Goyal , Mitesh M. Khapra , Anoop Kunchukuttan , Pratyush Kumar

Polarization is defined as divisive opinions held by two or more groups on substantive issues. As the world's third-largest democracy, Indonesia faces growing concerns about the interplay between political polarization and online toxicity,…

Computation and Language · Computer Science 2025-03-04 Lucky Susanto , Musa Wijanarko , Prasetia Pratama , Zilu Tang , Fariz Akyas , Traci Hong , Ika Idris , Alham Aji , Derry Wijaya

This tutorial (https://tum-nlp.github.io/low-resource-tutorial) is designed for NLP practitioners, researchers, and developers working with multilingual and low-resource languages who seek to create more equitable and socially impactful…

Computation and Language · Computer Science 2025-12-17 Ekaterina Artemova , Laurie Burchell , Daryna Dementieva , Shu Okabe , Mariya Shmatova , Pedro Ortiz Suarez

This paper presents Yankari, a large-scale monolingual dataset for the Yoruba language, aimed at addressing the critical gap in Natural Language Processing (NLP) resources for this important West African language. Despite being spoken by…

Computation and Language · Computer Science 2025-08-08 Maro Akpobi

A recurring challenge of crowdsourcing NLP datasets at scale is that human writers often rely on repetitive patterns when crafting examples, leading to a lack of linguistic diversity. We introduce a novel approach for dataset creation based…

Computation and Language · Computer Science 2022-11-16 Alisa Liu , Swabha Swayamdipta , Noah A. Smith , Yejin Choi