English
Related papers

Related papers: Advancing Uto-Aztecan Language Technologies: A Cas…

200 papers

Historically, researchers and consumers have noticed a decrease in quality when applying NLP tools to minority variants of languages (i.e. Puerto Rican Spanish or Swiss German), but studies exploring this have been limited to a select few…

Computation and Language · Computer Science 2023-10-24 Anjali Kantharuban , Ivan Vulić , Anna Korhonen

Recent advances in large language models (LLMs) have demonstrated remarkable capabilities on widely benchmarked high-resource languages. However, linguistic nuances of under-resourced languages remain unexplored. We introduce Batayan, a…

We present the first systematic study of machine translation for Chakma, an endangered and extremely low-resource Indo-Aryan language, with the goal of supporting language access and preservation. We introduce a new Chakma-Bangla parallel…

Computation and Language · Computer Science 2026-01-09 Aunabil Chakma , Aditya Chakma , Masum Hasan , Soham Khisa , Chumui Tripura , Rifat Shahriyar

Large Language Models (LLMs) demonstrate exceptional zero-shot capabilities in various NLP tasks, significantly enhancing user experience and efficiency. However, this advantage is primarily limited to resource-rich languages. For the…

Computation and Language · Computer Science 2025-09-23 Wenhao Zhuang , Yuan Sun

Despite excellent results on benchmarks over a small subset of languages, large language models struggle to process text from languages situated in `lower-resource' scenarios such as dialects/sociolects (national or social varieties of a…

Computation and Language · Computer Science 2024-09-20 Aditya Joshi , Diptesh Kanojia , Heather Lent , Hour Kaing , Haiyue Song

Natural language processing (NLP) has a significant impact on society via technologies such as machine translation and search engines. Despite its success, NLP technology is only widely available for high-resource languages such as English…

Lakota, a critically endangered language of the Sioux people in North America, faces significant challenges due to declining fluency among younger generations. This paper introduces LakotaBERT, the first large language model (LLM) tailored…

Computation and Language · Computer Science 2025-03-25 Kanishka Parankusham , Rodrigue Rizk , KC Santosh

The preservation of under-resourced languages requires digital tools and resources shaped by and for their speakers. We present the first dedicated ASR resources for Puno Quechua (ISO 639-3: qxp): (1) the largest speech corpus for any…

Computation and Language · Computer Science 2026-05-28 Elwin Huaman , Adrian Gamarra Lafuente , Johanna Cordova , Anna Korhonen

Preserving linguistic diversity is necessary as every language offers a distinct perspective on the world. There have been numerous global initiatives to preserve endangered languages through documentation. This paper is a part of a project…

Computation and Language · Computer Science 2025-10-28 Ambalika Guha , Sajal Saha , Debanjan Ballav , Soumi Mitra , Hritwick Chakraborty

Recent development of large-scale pre-trained language models (PLM) have significantly improved the capability of models in various NLP tasks, in terms of performance after task-specific fine-tuning and zero-shot / few-shot learning.…

Computation and Language · Computer Science 2022-04-21 Chenguang Zhu , Michael Zeng

Low-resource languages serve as invaluable repositories of human history, embodying cultural evolution and intellectual diversity. Despite their significance, these languages face critical challenges, including data scarcity and…

We develop machine translation and speech synthesis systems to complement the efforts of revitalizing Judeo-Spanish, the exiled language of Sephardic Jews, which survived for centuries, but now faces the threat of extinction in the digital…

Computation and Language · Computer Science 2022-06-01 Alp Öktem , Rodolfo Zevallos , Yasmin Moslem , Güneş Öztürk , Karen Şarhon

In rural regions of several developing countries, access to quality healthcare, medical infrastructure, and professional diagnosis is largely unavailable. Many of these regions are gradually gaining access to internet infrastructure,…

Computation and Language · Computer Science 2021-06-03 Vishal Vinod , Susmit Agrawal , Vipul Gaurav , Pallavi R , Savita Choudhary

Advances in large language models (LLMs) have empowered a variety of applications. However, there is still a significant gap in research when it comes to understanding and enhancing the capabilities of LLMs in the field of mental health. In…

Computation and Language · Computer Science 2024-01-30 Xuhai Xu , Bingsheng Yao , Yuanzhe Dong , Saadia Gabriel , Hong Yu , James Hendler , Marzyeh Ghassemi , Anind K. Dey , Dakuo Wang

This study investigates the potential of Large Language Models (LLMs), particularly GPT-4o, for Optical Character Recognition (OCR) in low-resource scripts such as Urdu, Albanian, and Tajik, with English serving as a benchmark. Using a…

Machine Learning · Computer Science 2024-12-23 Muhammad Abdullah Sohail , Salaar Masood , Hamza Iqbal

Propinquity between Australian Indigenous communities' social structures and ICT purposed for cultural preservation is a modern area of research; historically hindered by the "digital divide" thus limiting plentiful literature and existing…

Computers and Society · Computer Science 2016-06-07 Sarah Van Der Meer , Stephen Smith , Vincent Pang

How can large language models (LLMs) process and translate endangered languages? Many languages lack a large corpus to train a decent LLM; therefore existing LLMs rarely perform well in unseen, endangered languages. On the contrary, we…

Computation and Language · Computer Science 2024-11-13 Kexun Zhang , Yee Man Choi , Zhenqiao Song , Taiqi He , William Yang Wang , Lei Li

Natural language processing (NLP) practitioners are leveraging large language models (LLM) to create structured datasets from semi-structured and unstructured data sources such as patents, papers, and theses, without having domain-specific…

Computation and Language · Computer Science 2024-03-26 Jesse Atuhurra , Seiveright Cargill Dujohn , Hidetaka Kamigaito , Hiroyuki Shindo , Taro Watanabe

The vast majority of the world's languages, particularly creoles like Nagamese, remain severely under-resourced in Natural Language Processing (NLP), creating a significant barrier to their representation in digital technology. This paper…

Computation and Language · Computer Science 2025-12-16 Agniva Maiti , Manya Pandey , Murari Mandal

We are 600 million Spanish speakers. We launched the #Somos600M Project because the diversity of the languages from LATAM, the Caribbean and Spain needs to be represented in Artificial Intelligence (AI) systems. Despite being the 7.5% of…

Computation and Language · Computer Science 2024-07-26 María Grandury