中文
相关论文

相关论文: Survey on Publicly Available Sinhala Natural Langu…

200 篇论文

This paper introduces the development of the first open conversational speech dataset for the Isan language, the most widely spoken regional dialect in Thailand. Unlike existing speech corpora that are primarily based on read or scripted…

计算与语言 · 计算机科学 2025-12-05 Adisai Na-Thalang , Chanakan Wittayasakpan , Kritsadha Phatcharoen , Supakit Buakaw

Deep neural networks and huge language models are becoming omnipresent in natural language applications. As they are known for requiring large amounts of training data, there is a growing body of work to improve the performance in…

计算与语言 · 计算机科学 2021-04-12 Michael A. Hedderich , Lukas Lange , Heike Adel , Jannik Strötgen , Dietrich Klakow

Language sample analysis (LSA) is a process that complements standardized psychometric tests for diagnosing, for example, developmental language disorder (DLD) in children. However, its labour-intensive nature has limited its use in…

计算与语言 · 计算机科学 2026-03-31 Anja Ryser , Yingqiang Gao , Sarah Ebling

Speech large language models (SLLMs) built on speech encoders, adapters, and LLMs demonstrate remarkable multitask understanding performance in high-resource languages such as English and Chinese. However, their effectiveness substantially…

声音 · 计算机科学 2026-04-21 Mingchen Shao , Bingshen Mu , Chengyou Wang , Hai Li , Ying Yan , Zhonghua Fu , Lei Xie

Enabling farmers to access accurate agriculture-related information in their native languages in a timely manner is crucial for the success of the agriculture field. Publicly available general-purpose Large Language Models (LLMs) typically…

Most low-resource languages do not have the necessary resources to create even a substantial monolingual corpus. These languages may often be found in government proceedings but mainly in Portable Document Format (PDF) that contains legacy…

计算与语言 · 计算机科学 2022-12-19 Charangan Vasantharajan , Laksika Tharmalingam , Uthayasanker Thayasivam

Indigenous languages of the American continent are highly diverse. However, they have received little attention from the technological perspective. In this paper, we review the research, the digital resources and the available NLP systems…

计算与语言 · 计算机科学 2018-06-13 Manuel Mager , Ximena Gutierrez-Vasques , Gerardo Sierra , Ivan Meza

The rise of large language models (LLMs) has led many researchers to consider their usage for scientific work. Some have found benefits using LLMs to augment or automate aspects of their research pipeline, while others have urged caution…

计算与语言 · 计算机科学 2024-11-11 Zhehui Liao , Maria Antoniak , Inyoung Cheong , Evie Yu-Yen Cheng , Ai-Heng Lee , Kyle Lo , Joseph Chee Chang , Amy X. Zhang

Despite much progress in recent years, the vast majority of work in natural language processing (NLP) is on standard languages with many speakers. In this work, we instead focus on low-resource languages and in particular non-standardized…

计算与语言 · 计算机科学 2023-04-20 Verena Blaschke , Hinrich Schütze , Barbara Plank

Natural language processing is a branch of computer science that combines artificial intelligence with linguistics. It aims to analyze a language element such as writing or speaking with software and convert it into information. Considering…

计算与语言 · 计算机科学 2021-01-28 Kadir Tohma , Yakup Kutlu

As large language models (LLMs) expand multilingual capabilities, questions remain about the equity of their performance across languages. While many communities stand to benefit from AI systems, the dominance of English in training data…

计算与语言 · 计算机科学 2025-09-30 Sophie Jaffer , Simeon Sayer

Out of vocabulary (OOV) is a problem in the context of Machine Translation (MT) in low-resourced languages. When source and/or target languages are morphologically rich, it becomes even worse. Bilingual list integration is an approach to…

计算与语言 · 计算机科学 2021-02-04 Aloka Fernando , Surangika Ranathunga , Gihan Dias

India's linguistic landscape, spanning 22 scheduled languages and hundreds of marginalized dialects, has driven rapid growth in NLP datasets, benchmarks, and pretrained models. However, no dedicated survey consolidates resources developed…

计算与语言 · 计算机科学 2026-04-21 Raghvendra Kumar , Devankar Raj , Sriparna Saha

Indigenous African languages are categorized as under-served in Natural Language Processing. They therefore experience poor digital inclusivity and information access. The processing challenge with such languages has been how to use machine…

计算与语言 · 计算机科学 2025-01-17 Barack Wanjawa , Lilian Wanzare , Florence Indede , Owen McOnyango , Edward Ombui , Lawrence Muchemi

Unsupervised translation has reached impressive performance on resource-rich language pairs such as English-French and English-German. However, early studies have shown that in more realistic settings involving low-resource, rare languages,…

计算与语言 · 计算机科学 2021-03-15 Xavier Garcia , Aditya Siddhant , Orhan Firat , Ankur P. Parikh

State-of-the-art natural language processing (NLP) models are trained on massive training corpora, and report a superlative performance on evaluation datasets. This survey delves into an important attribute of these datasets: the dialect of…

计算与语言 · 计算机科学 2024-12-10 Aditya Joshi , Raj Dabre , Diptesh Kanojia , Zhuang Li , Haolan Zhan , Gholamreza Haffari , Doris Dippold

We report findings of the TSAR-2022 shared task on multilingual lexical simplification, organized as part of the Workshop on Text Simplification, Accessibility, and Readability TSAR-2022 held in conjunction with EMNLP 2022. The task called…

计算与语言 · 计算机科学 2023-02-07 Horacio Saggion , Sanja Štajner , Daniel Ferrés , Kim Cheng Sheang , Matthew Shardlow , Kai North , Marcos Zampieri

Recent advancement in large language models (LLMs) has offered a strong potential for natural language systems to process informal language. A representative form of informal language is slang, used commonly in daily conversations and…

计算与语言 · 计算机科学 2024-04-16 Zhewei Sun , Qian Hu , Rahul Gupta , Richard Zemel , Yang Xu

NLP research is impeded by a lack of resources and awareness of the challenges presented by underrepresented languages and dialects. Focusing on the languages spoken in Indonesia, the second most linguistically diverse and the fourth most…