English
Related papers

Related papers: NusaWrites: Constructing High-Quality Corpora for …

200 papers

Natural language processing (NLP) has grown significantly since the advent of the Transformer architecture. Transformers have given birth to pre-trained large language models (PLMs). There has been tremendous improvement in the performance…

Computation and Language · Computer Science 2024-07-16 Jesse Atuhurra , Hidetaka Kamigaito

With the growing presence of multilingual users on social media, detecting abusive language in code-mixed text has become increasingly challenging. Code-mixed communication, where users seamlessly switch between English and their native…

Computation and Language · Computer Science 2025-05-01 Manish Pandey , Nageshwar Prasad Yadav , Mokshada Adduru , Sawan Rai

Dyslexia in adults remains an under-researched and under-served area, particularly in non-English-speaking contexts, despite its significant impact on personal and professional lives. This work addresses that gap by focusing on Sinhala, a…

Computation and Language · Computer Science 2025-10-07 Peshala Perera , Deshan Sumanathilaka

This paper describes a web-based corpus of global language use with a focus on how this corpus can be used for data-driven language mapping. First, the corpus provides a representation of where national varieties of major languages are used…

Computation and Language · Computer Science 2020-04-03 Jonathan Dunn

Compared to English, the amount of labeled data for Indonesian text classification tasks is very small. Recently developed multilingual language models have shown its ability to create multilingual representations effectively. This paper…

Computation and Language · Computer Science 2020-09-15 Ilham Firdausi Putra , Ayu Purwarianti

Stereotype repositories are critical to assess generative AI model safety, but currently lack adequate global coverage. It is imperative to prioritize targeted expansion, strategically addressing existing deficits, over merely increasing…

Computation and Language · Computer Science 2026-02-27 Aishwarya Verma , Laud Ammah , Olivia Nercy Ndlovu Lucas , Andrew Zaldivar , Vinodkumar Prabhakaran , Sunipa Dev

Performance of NLP systems is typically evaluated by collecting a large-scale dataset by means of crowd-sourcing to train a data-driven model and evaluate it on a held-out portion of the data. This approach has been shown to suffer from…

Computation and Language · Computer Science 2024-08-12 Viktor Schlegel , Goran Nenadic , Riza Batista-Navarro

Multilingual NLP often relies on dataset counts from centralized catalogues to characterize which languages are resource-rich or resource-poor. However, these catalogues record only one layer of dataset visibility: what has been registered…

Computation and Language · Computer Science 2026-05-19 Zhiyin Tan , Changxu Duan

In this work, we present BanglaParaphrase, a high-quality synthetic Bangla Paraphrase dataset curated by a novel filtering pipeline. We aim to take a step towards alleviating the low resource status of the Bangla language in the NLP domain…

Computation and Language · Computer Science 2022-10-12 Ajwad Akil , Najrin Sultana , Abhik Bhattacharjee , Rifat Shahriyar

Generative Artificial Intelligence (GenAI), particularly Large Language Models (LLMs), has significantly advanced Natural Language Processing (NLP) tasks, such as Named Entity Recognition (NER), which involves identifying entities like…

Computation and Language · Computer Science 2025-03-14 Sameer Neupane , Jeevan Chapagain , Nobal B. Niraula , Diwa Koirala

This preprint describes work in progress on LR-Sum, a new permissively-licensed dataset created with the goal of enabling further research in automatic summarization for less-resourced languages. LR-Sum contains human-written summaries for…

Computation and Language · Computer Science 2023-10-30 Chester Palen-Michel , Constantine Lignos

Parallel corpora play an important role in training machine translation (MT) models, particularly for low-resource languages where high-quality bilingual data is scarce. This review provides a comprehensive overview of available parallel…

Computation and Language · Computer Science 2025-04-23 Rahul Raja , Arpita Vats

Large text corpora are increasingly important for a wide variety of Natural Language Processing (NLP) tasks, and automatic language identification (LangID) is a core technology needed to collect such datasets in a multilingual context.…

Computation and Language · Computer Science 2020-10-30 Isaac Caswell , Theresa Breiner , Daan van Esch , Ankur Bapna

[Abridged Abstract] Recent technological advances underscore labor market dynamics, yielding significant consequences for employment prospects and increasing job vacancy data across platforms and languages. Aggregating such data holds…

Computation and Language · Computer Science 2024-05-01 Mike Zhang

Indonesian language is spoken by almost 200 million people and is the 10th most spoken language in the world, but it is under-represented in NLP (Natural Language Processing) research. A sparsity of language resources has hampered previous…

Computation and Language · Computer Science 2024-10-28 Mukhlish Fuadi , Adhi Dharma Wibawa , Surya Sumpeno

Massively multilingual neural machine translation (MMNMT) has been proven to enhance the translation quality of low-resource languages. In this paper, we empirically investigate the translation robustness of Indonesian-Chinese translation…

Computation and Language · Computer Science 2024-05-14 Supryadi , Leiyu Pan , Deyi Xiong

Sign languages serve as essential communication systems for individuals with hearing and speech impairments. However, digital linguistic dataset resources for underrepresented sign languages, such as Nepali Sign Language (NSL), remain…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Birat Poudel , Satyam Ghimire , Sijan Bhattarai , Saurav Bhandari , Suramya Sharma Dahal

Sentence-level embedding is essential for various tasks that require understanding natural language. Many studies have explored such embeddings for high-resource languages like English. However, low-resource languages like Bengali (a…

Computation and Language · Computer Science 2024-11-26 Muhammad Rafsan Kabir , Md. Mohibur Rahman Nabil , Mohammad Ashrafuzzaman Khan

The availability of corpora is a major factor in building natural language processing applications. However, the costs of acquiring corpora can prevent some researchers from going further in their endeavours. The ease of access to freely…

Computation and Language · Computer Science 2017-02-28 Wajdi Zaghouani

Unlike mainstream languages (such as English and French), low-resource languages often suffer from a lack of expert-annotated corpora and benchmark resources that make it hard to apply state-of-the-art techniques directly. In this paper, we…

Computation and Language · Computer Science 2019-07-04 Jan Christian Blaise Cruz , Charibeth Cheng
‹ Prev 1 8 9 10 Next ›