English
Related papers

Related papers: IndoLEM and IndoBERT: A Benchmark Dataset and Pre-…

200 papers

Although Indonesian is known to be the fourth most frequently used language over the internet, the research progress on this language in the natural language processing (NLP) is slow-moving due to a lack of available resources. In response,…

Over 200 million people speak Indonesian, yet the language remains significantly underrepresented in preference-based research for large language models (LLMs). Most existing multilingual datasets are derived from English translations,…

Computation and Language · Computer Science 2025-11-13 Vanessa Rebecca Wiyono , David Anugraha , Ayu Purwarianti , Genta Indra Winata

Indonesia's linguistic landscape is remarkably diverse, encompassing over 700 languages and dialects, making it one of the world's most linguistically rich nations. This diversity, coupled with the widespread practice of code-switching and…

Computation and Language · Computer Science 2024-03-05 Wilson Wongso , David Samuel Setiawan , Steven Limcorn , Ananto Joyoadikusumo

Natural language generation (NLG) benchmarks provide an important avenue to measure progress and develop better NLG systems. Unfortunately, the lack of publicly available NLG benchmarks for low-resource languages poses a challenging barrier…

Automatic text summarization is generally considered as a challenging task in the NLP community. One of the challenges is the publicly available and large dataset that is relatively rare and difficult to construct. The problem is even worse…

Computation and Language · Computer Science 2019-03-21 Kemal Kurniawan , Samuel Louvan

NLP research is impeded by a lack of resources and awareness of the challenges presented by underrepresented languages and dialects. Focusing on the languages spoken in Indonesia, the second most linguistically diverse and the fourth most…

We present IndoNLI, the first human-elicited NLI dataset for Indonesian. We adapt the data collection protocol for MNLI and collect nearly 18K sentence pairs annotated by crowd workers and experts. The expert-annotated data is used…

Computation and Language · Computer Science 2022-03-30 Rahmad Mahendra , Alham Fikri Aji , Samuel Louvan , Fahrurrozi Rahman , Clara Vania

As one of the world's most populous countries, with 700 languages spoken, Indonesia is behind in terms of NLP progress. We introduce LoraxBench, a benchmark that focuses on low-resource languages of Indonesia and covers 6 diverse tasks:…

Computation and Language · Computer Science 2025-08-19 Alham Fikri Aji , Trevor Cohn

Natural language processing (NLP) has a significant impact on society via technologies such as machine translation and search engines. Despite its success, NLP technology is only widely available for high-resource languages such as English…

Understanding emotions in the Indonesian language is essential for improving customer experiences in e-commerce. This study focuses on enhancing the accuracy of emotion classification in Indonesian by leveraging advanced language models,…

Computation and Language · Computer Science 2025-09-19 William Christian , Daniel Adamlu , Adrian Yu , Derwin Suhartono

BERT and IndoBERT have achieved impressive performance in several NLP tasks. There has been several investigation on its adaption in specialized domains especially for English language. We focus on financial domain and Indonesian language,…

Computation and Language · Computer Science 2023-10-17 Ni Putu Intan Maharani , Yoga Yustiawan , Fauzy Caesar Rochim , Ayu Purwarianti

Determining whether a piece of text is relevant to a given topic is a fundamental task in natural language processing, yet it remains largely unexplored for Bahasa Indonesia. Unlike sentiment analysis or named entity recognition, relevancy…

Indonesian language is spoken by almost 200 million people and is the 10th most spoken language in the world, but it is under-represented in NLP (Natural Language Processing) research. A sparsity of language resources has hampered previous…

Computation and Language · Computer Science 2024-10-28 Mukhlish Fuadi , Adhi Dharma Wibawa , Surya Sumpeno

We present IndoBERTweet, the first large-scale pretrained model for Indonesian Twitter that is trained by extending a monolingually-trained Indonesian BERT model with additive domain-specific vocabulary. We focus in particular on efficient…

Computation and Language · Computer Science 2021-09-13 Fajri Koto , Jey Han Lau , Timothy Baldwin

In this work, we introduce BanglaBERT, a BERT-based Natural Language Understanding (NLU) model pretrained in Bangla, a widely spoken yet low-resource language in the NLP literature. To pretrain BanglaBERT, we collect 27.5 GB of Bangla…

Computation and Language · Computer Science 2022-05-11 Abhik Bhattacharjee , Tahmid Hasan , Wasi Uddin Ahmad , Kazi Samin , Md Saiful Islam , Anindya Iqbal , M. Sohel Rahman , Rifat Shahriyar

Indonesian, spoken by over 200 million people, remains underserved in multimodal emotion recognition research despite its dominant presence on Southeast Asian social media platforms. We introduce IndoMER, the first multimodal emotion…

Machine Learning · Computer Science 2026-02-11 Xueming Yan , Boyan Xu , Yaochu Jin , Lixian Xiao , Wenlong Ye , Runyang Cai , Zeqi Zheng , Jingfa Liu , Aimin Yang , Yongduan Song

Hate speech poses a significant threat to social harmony. Over the past two years, Indonesia has seen a ten-fold increase in the online hate speech ratio, underscoring the urgent need for effective detection mechanisms. However, progress is…

Computation and Language · Computer Science 2025-06-13 Lucky Susanto , Musa Izzanardi Wijanarko , Prasetia Anugrah Pratama , Traci Hong , Ika Idris , Alham Fikri Aji , Derry Wijaya

Although large language models (LLMs) are often pre-trained on large-scale multilingual texts, their reasoning abilities and real-world knowledge are mainly evaluated based on English datasets. Assessing LLM capabilities beyond English is…

Computation and Language · Computer Science 2023-10-24 Fajri Koto , Nurul Aisyah , Haonan Li , Timothy Baldwin

Multimodal learning on video and text has seen significant progress, particularly in tasks like text-to-video retrieval, video-to-text retrieval, and video captioning. However, most existing methods and datasets focus exclusively on…

Multimedia · Computer Science 2025-07-15 Willy Fitra Hendria

At the center of the underlying issues that halt Indonesian natural language processing (NLP) research advancement, we find data scarcity. Resources in Indonesian languages, especially the local ones, are extremely scarce and…

‹ Prev 1 2 3 10 Next ›