中文
相关论文

相关论文: On Language Models for Creoles

200 篇论文

Large Language Models (LLMs) have been reported to have strong performance on natural language processing tasks. However, performance metrics such as accuracy do not measure the quality of the model in terms of its ability to robustly…

机器学习 · 计算机科学 2023-06-02 Emanuele La Malfa , Matthew Wicker , Marta Kwiatkowska

Low-resource African languages pose unique challenges for natural language processing (NLP) tasks, including natural language generation (NLG). In this paper, we develop Cheetah, a massively multilingual NLG language model for African…

计算与语言 · 计算机科学 2024-01-11 Ife Adebara , AbdelRahim Elmadany , Muhammad Abdul-Mageed

Social bias in language models can potentially exacerbate social inequalities. Despite it having garnered wide attention, most research focuses on English data. In a low-resource scenario, the models often perform worse due to insufficient…

计算与语言 · 计算机科学 2025-07-15 Ej Zhou , Weiming Lu

This paper investigates the impact of corpus creation decisions on large multi-lingual geographic web corpora. Beginning with a 427 billion word corpus derived from the Common Crawl, three methods are used to improve the quality of…

计算与语言 · 计算机科学 2024-03-14 Jonathan Dunn

Neural language models have achieved state-of-the-art performances on many NLP tasks, and recently have been shown to learn a number of hierarchically-sensitive syntactic dependencies between individual words. However, equally important for…

计算与语言 · 计算机科学 2019-09-11 Aixiu An , Peng Qian , Ethan Wilcox , Roger Levy

In recent years, the field of Natural Language Generation (NLG) has been boosted by the recent advances in deep learning technologies. Nonetheless, these new data-intensive methods introduce language-dependent disparities in NLG as the main…

Multilingual pretrained language models (mPLMs) acquire valuable, generalizable linguistic information during pretraining and have advanced the state of the art on task-specific finetuning. To date, only ~31 out of ~2,000 African languages…

计算与语言 · 计算机科学 2023-05-30 Ife Adebara , AbdelRahim Elmadany , Muhammad Abdul-Mageed , Alcides Alcoba Inciarte

Multilingual Retrieval-Augmented Generation (mRAG) systems enable language models to answer knowledge-intensive queries with citation-supported responses across languages. While such systems have been proposed, an open questions is whether…

计算与语言 · 计算机科学 2025-10-03 Dayeon Ki , Marine Carpuat , Paul McNamee , Daniel Khashabi , Eugene Yang , Dawn Lawrie , Kevin Duh

Bias studies on multilingual models confirm the presence of gender-related stereotypes in masked models processing languages with high NLP resources. We expand on this line of research by introducing Filipino CrowS-Pairs and Filipino…

计算与语言 · 计算机科学 2025-04-29 Lance Calvin Lim Gamboa , Mark Lee

Building natural language processing systems for non standardized and low resource languages is a difficult challenge. The recent success of large-scale multilingual pretrained language models provides new modeling tools to tackle this. In…

计算与语言 · 计算机科学 2020-05-04 Benjamin Muller , Benoit Sagot , Djamé Seddah

Cross-lingual transfer has become a crucial aspect of multilingual NLP, as it allows for models trained on resource-rich languages to be applied to low-resource languages more effectively. Recently massively multilingual pre-trained…

计算与语言 · 计算机科学 2025-05-21 Ajitesh Bankula , Praney Bankula

Norwegian, spoken by only 5 million population, is under-representative within the most impressive breakthroughs in NLP tasks. To the best of our knowledge, there has not yet been a comprehensive evaluation of the existing language models…

计算与语言 · 计算机科学 2024-10-02 Peng Liu , Lemei Zhang , Terje Farup , Even W. Lauvrak , Jon Espen Ingvaldsen , Simen Eide , Jon Atle Gulla , Zhirong Yang

In this study, we investigate whether non-English-centric LLMs, despite their strong performance, `think' in their respective dominant language: more precisely, `think' refers to how the representations of intermediate layers, when…

计算与语言 · 计算机科学 2024-08-21 Chengzhi Zhong , Fei Cheng , Qianying Liu , Junfeng Jiang , Zhen Wan , Chenhui Chu , Yugo Murawaki , Sadao Kurohashi

Word embeddings and pre-trained language models allow to build rich representations of text and have enabled improvements across most NLP tasks. Unfortunately they are very expensive to train, and many small companies and research groups…

计算与语言 · 计算机科学 2020-04-03 Rodrigo Agerri , Iñaki San Vicente , Jon Ander Campos , Ander Barrena , Xabier Saralegi , Aitor Soroa , Eneko Agirre

The Nagamese language, a.k.a Naga Pidgin, is an Assamese-lexified creole language developed primarily as a means of communication in trade between the people from Nagaland and people from Assam in the north-east India. Substantial amount of…

计算与语言 · 计算机科学 2025-12-19 Ekha Morang , Surhoni A. Ngullie , Sashienla Longkumer , Teisovi Angami

Most current large language models (LLMs) support a wide variety of languages in addition to English, including high-resource languages (e.g. German, Chinese, French), as well as low-resource ones (e.g. Swahili, Telugu). In addition they…

计算与语言 · 计算机科学 2025-11-10 Jan-Thorsten Peter , David Vilar , Tobias Domhan , Dan Malkin , Markus Freitag

In cross-lingual transfer, NLP models over one or more source languages are applied to a low-resource target language. While most prior work has used a single source model or a few carefully selected models, here we consider a `massive'…

计算与语言 · 计算机科学 2019-06-06 Afshin Rahimi , Yuan Li , Trevor Cohn

The dominance of large multilingual foundation models has widened linguistic inequalities in Natural Language Processing (NLP), often leaving low-resource languages underrepresented. This paper introduces LilMoo, a 0.6-billion-parameter…

计算与语言 · 计算机科学 2026-03-05 Shiza Fatimah , Aniket Sen , Sophia Falk , Florian Mai , Lucie Flek , Nicholas Kluge Corrêa

Spelling normalization for low resource languages is a challenging task because the patterns are hard to predict and large corpora are usually required to collect enough examples. This work shows a comparison of a neural model and character…

计算与语言 · 计算机科学 2020-10-21 Yiyuan Li , Antonios Anastasopoulos , Alan W Black

Large language models (LLMs) have demonstrated impressive capabilities across various natural language processing (NLP) tasks in recent years. However, their susceptibility to jailbreaks and perturbations necessitates additional…

计算与语言 · 计算机科学 2025-06-10 Maciej Chrabąszcz , Katarzyna Lorenc , Karolina Seweryn
‹ 上一页 1 8 9 10 下一页 ›