English
Related papers

Related papers: FormosanBench: Benchmarking Low-Resource Austrones…

200 papers

We introduce BURMESE-SAN, the first holistic benchmark that systematically evaluates large language models (LLMs) for Burmese across three core NLP competencies: understanding (NLU), reasoning (NLR), and generation (NLG). BURMESE-SAN…

Computation and Language · Computer Science 2026-05-25 Thura Aung , Jann Railey Montalan , Jian Gang Ngui , Peerat Limkonchotiwat

Recent advances in large language models (LLMs) have demonstrated remarkable capabilities on widely benchmarked high-resource languages. However, linguistic nuances of under-resourced languages remain unexplored. We introduce Batayan, a…

Large Language Models (LLMs) have demonstrated remarkable success across a wide range of tasks and domains. However, their performance in low-resource language translation, particularly when translating into these languages, remains…

This study explores the use of large language models (LLMs) for translating English into Mambai, a low-resource Austronesian language spoken in Timor-Leste, with approximately 200,000 native speakers. Leveraging a novel corpus derived from…

Computation and Language · Computer Science 2025-01-28 Raphaël Merx , Aso Mahmudi , Katrina Langford , Leo Alberto de Araujo , Ekaterina Vylomova

Despite the widespread adoption of Large language models (LLMs), their remarkable capabilities remain limited to a few high-resource languages. Additionally, many low-resource languages (\eg African languages) are often evaluated only on…

Large language models (LLMs) have driven substantial advances in speech language models (SpeechLMs), yielding strong performance in automatic speech recognition (ASR) under high-resource conditions. However, existing benchmarks…

Computation and Language · Computer Science 2026-03-23 Jianan Chen , Xiaoxue Gao , Tatsuya Kawahara , Nancy F. Chen

Large Language Models (LLMs) have shown remarkable abilities across various tasks, yet their development has predominantly centered on high-resource languages like English and Chinese, leaving low-resource languages underserved. To address…

Computation and Language · Computer Science 2024-07-30 Wenxuan Zhang , Hou Pong Chan , Yiran Zhao , Mahani Aljunied , Jianyu Wang , Chaoqun Liu , Yue Deng , Zhiqiang Hu , Weiwen Xu , Yew Ken Chia , Xin Li , Lidong Bing

Large Language Models (LLMs) have demonstrated remarkable performance across various Natural Language Processing (NLP) tasks, largely due to their generalisability and ability to perform tasks without additional training. However, their…

Computation and Language · Computer Science 2025-08-15 Kurt Micallef , Claudia Borg

Rapid developments of large language models have revolutionized many NLP tasks for English data. Unfortunately, the models and their evaluations for low-resource languages are being overlooked, especially for languages in South Asia.…

Computation and Language · Computer Science 2025-09-16 Sampoorna Poria , Xiaolei Huang

Australian Aboriginal languages are of significant cultural and linguistic value but remain severely underrepresented in modern speech AI systems. While state-of-the-art speech foundation models and automatic speech recognition excel in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Ting Dang , Trini Manoj Jeyaseelan , Eliathamby Ambikairajah , Vidhyasaharan Sethu

The rapid advancement of large language models (LLMs) has not been matched by their evaluation in low-resource languages, especially Southeast Asian languages like Lao. To fill this gap, we introduce \textbf{LaoBench}, the first…

Computation and Language · Computer Science 2026-04-16 Jian Gao , Richeng Xuan , Zhaolu Kang , Dingshi Liao , Wenxin Huang , Zongmou Huang , Yangdi Xu , Bowen Qin , Zheqi He , Xi Yang , Changjin Li , Yonghua Lin

The advancement of large language models (LLMs) has predominantly focused on high-resource languages, leaving low-resource languages, such as those in the Finno-Ugric family, significantly underrepresented. This paper addresses this gap by…

Computation and Language · Computer Science 2025-05-06 Taido Purason , Hele-Andra Kuulmets , Mark Fishel

The advent of Multilingual Language Models (MLLMs) and Large Language Models has spawned innovation in many areas of natural language processing. Despite the exciting potential of this technology, its impact on developing high-quality…

Computation and Language · Computer Science 2024-03-06 Séamus Lankford , Haithem Afli , Andy Way

This paper provides language identification models for low- and under-resourced languages in the Pacific region with a focus on previously unavailable Austronesian languages. Accurate language identification is an important part of…

Computation and Language · Computer Science 2022-06-10 Jonathan Dunn , Wikke Nijhof

Despite dramatic recent progress in NLP, it is still a major challenge to apply Large Language Models (LLM) to low-resource languages. This is made visible in benchmarks such as Cross-Lingual Natural Language Inference (XNLI), a key task…

Computation and Language · Computer Science 2025-04-15 Aung Kyaw Htet , Mark Dras

Despite the remarkable achievements of large language models (LLMs) in various tasks, there remains a linguistic bias that favors high-resource languages, such as English, often at the expense of low-resource and regional languages. To…

Automatic speech recognition systems have undoubtedly advanced with the integration of multilingual and multitask models such as Whisper, which have shown a promising ability to understand and process speech across a wide range of…

Computation and Language · Computer Science 2025-04-14 Xabier de Zuazo , Eva Navas , Ibon Saratxaga , Inma Hernáez Rioja

Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-resource languages, reaching state-of-the-art performance in various tasks. However, their applicability is still less explored in low-resource…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-08 Seraphina Fong , Marco Matassoni , Alessio Brutti

Low-resource languages pose persistent challenges for Natural Language Processing tasks such as lemmatization and part-of-speech (POS) tagging. This paper investigates the capacity of recent large language models (LLMs), including GPT-4…

Computation and Language · Computer Science 2026-02-18 Chahan Vidal-Gorène , Bastien Kindt , Florian Cafiero

Large Language Models (LLMs) have demonstrated exceptional promise in translation tasks for high-resource languages. However, their performance in low-resource languages is limited by the scarcity of both parallel and monolingual corpora,…

Computation and Language · Computer Science 2024-10-11 William Tan , Kevin Zhu
‹ Prev 1 2 3 10 Next ›