中文
相关论文

相关论文: Advancing Uto-Aztecan Language Technologies: A Cas…

200 篇论文

Historically, researchers and consumers have noticed a decrease in quality when applying NLP tools to minority variants of languages (i.e. Puerto Rican Spanish or Swiss German), but studies exploring this have been limited to a select few…

计算与语言 · 计算机科学 2023-10-24 Anjali Kantharuban , Ivan Vulić , Anna Korhonen

Recent advances in large language models (LLMs) have demonstrated remarkable capabilities on widely benchmarked high-resource languages. However, linguistic nuances of under-resourced languages remain unexplored. We introduce Batayan, a…

We present the first systematic study of machine translation for Chakma, an endangered and extremely low-resource Indo-Aryan language, with the goal of supporting language access and preservation. We introduce a new Chakma-Bangla parallel…

计算与语言 · 计算机科学 2026-01-09 Aunabil Chakma , Aditya Chakma , Masum Hasan , Soham Khisa , Chumui Tripura , Rifat Shahriyar

Large Language Models (LLMs) demonstrate exceptional zero-shot capabilities in various NLP tasks, significantly enhancing user experience and efficiency. However, this advantage is primarily limited to resource-rich languages. For the…

计算与语言 · 计算机科学 2025-09-23 Wenhao Zhuang , Yuan Sun

Despite excellent results on benchmarks over a small subset of languages, large language models struggle to process text from languages situated in `lower-resource' scenarios such as dialects/sociolects (national or social varieties of a…

计算与语言 · 计算机科学 2024-09-20 Aditya Joshi , Diptesh Kanojia , Heather Lent , Hour Kaing , Haiyue Song

Natural language processing (NLP) has a significant impact on society via technologies such as machine translation and search engines. Despite its success, NLP technology is only widely available for high-resource languages such as English…

Lakota, a critically endangered language of the Sioux people in North America, faces significant challenges due to declining fluency among younger generations. This paper introduces LakotaBERT, the first large language model (LLM) tailored…

计算与语言 · 计算机科学 2025-03-25 Kanishka Parankusham , Rodrigue Rizk , KC Santosh

The preservation of under-resourced languages requires digital tools and resources shaped by and for their speakers. We present the first dedicated ASR resources for Puno Quechua (ISO 639-3: qxp): (1) the largest speech corpus for any…

计算与语言 · 计算机科学 2026-05-28 Elwin Huaman , Adrian Gamarra Lafuente , Johanna Cordova , Anna Korhonen

Preserving linguistic diversity is necessary as every language offers a distinct perspective on the world. There have been numerous global initiatives to preserve endangered languages through documentation. This paper is a part of a project…

计算与语言 · 计算机科学 2025-10-28 Ambalika Guha , Sajal Saha , Debanjan Ballav , Soumi Mitra , Hritwick Chakraborty

Recent development of large-scale pre-trained language models (PLM) have significantly improved the capability of models in various NLP tasks, in terms of performance after task-specific fine-tuning and zero-shot / few-shot learning.…

计算与语言 · 计算机科学 2022-04-21 Chenguang Zhu , Michael Zeng

Low-resource languages serve as invaluable repositories of human history, embodying cultural evolution and intellectual diversity. Despite their significance, these languages face critical challenges, including data scarcity and…

We develop machine translation and speech synthesis systems to complement the efforts of revitalizing Judeo-Spanish, the exiled language of Sephardic Jews, which survived for centuries, but now faces the threat of extinction in the digital…

计算与语言 · 计算机科学 2022-06-01 Alp Öktem , Rodolfo Zevallos , Yasmin Moslem , Güneş Öztürk , Karen Şarhon

In rural regions of several developing countries, access to quality healthcare, medical infrastructure, and professional diagnosis is largely unavailable. Many of these regions are gradually gaining access to internet infrastructure,…

计算与语言 · 计算机科学 2021-06-03 Vishal Vinod , Susmit Agrawal , Vipul Gaurav , Pallavi R , Savita Choudhary

Advances in large language models (LLMs) have empowered a variety of applications. However, there is still a significant gap in research when it comes to understanding and enhancing the capabilities of LLMs in the field of mental health. In…

计算与语言 · 计算机科学 2024-01-30 Xuhai Xu , Bingsheng Yao , Yuanzhe Dong , Saadia Gabriel , Hong Yu , James Hendler , Marzyeh Ghassemi , Anind K. Dey , Dakuo Wang

This study investigates the potential of Large Language Models (LLMs), particularly GPT-4o, for Optical Character Recognition (OCR) in low-resource scripts such as Urdu, Albanian, and Tajik, with English serving as a benchmark. Using a…

机器学习 · 计算机科学 2024-12-23 Muhammad Abdullah Sohail , Salaar Masood , Hamza Iqbal

Propinquity between Australian Indigenous communities' social structures and ICT purposed for cultural preservation is a modern area of research; historically hindered by the "digital divide" thus limiting plentiful literature and existing…

计算机与社会 · 计算机科学 2016-06-07 Sarah Van Der Meer , Stephen Smith , Vincent Pang

How can large language models (LLMs) process and translate endangered languages? Many languages lack a large corpus to train a decent LLM; therefore existing LLMs rarely perform well in unseen, endangered languages. On the contrary, we…

计算与语言 · 计算机科学 2024-11-13 Kexun Zhang , Yee Man Choi , Zhenqiao Song , Taiqi He , William Yang Wang , Lei Li

Natural language processing (NLP) practitioners are leveraging large language models (LLM) to create structured datasets from semi-structured and unstructured data sources such as patents, papers, and theses, without having domain-specific…

计算与语言 · 计算机科学 2024-03-26 Jesse Atuhurra , Seiveright Cargill Dujohn , Hidetaka Kamigaito , Hiroyuki Shindo , Taro Watanabe

The vast majority of the world's languages, particularly creoles like Nagamese, remain severely under-resourced in Natural Language Processing (NLP), creating a significant barrier to their representation in digital technology. This paper…

计算与语言 · 计算机科学 2025-12-16 Agniva Maiti , Manya Pandey , Murari Mandal

We are 600 million Spanish speakers. We launched the #Somos600M Project because the diversity of the languages from LATAM, the Caribbean and Spain needs to be represented in Artificial Intelligence (AI) systems. Despite being the 7.5% of…

计算与语言 · 计算机科学 2024-07-26 María Grandury