中文
相关论文

相关论文: Simple Additions, Substantial Gains: Expanding Scr…

200 篇论文

URIEL is a knowledge base offering geographical, phylogenetic, and typological vector representations for 7970 languages. It includes distance measures between these vectors for 4005 languages, which are accessible via the lang2vec tool.…

Existing linguistic knowledge bases such as URIEL+ provide valuable geographic, genetic and typological distances for cross-lingual transfer but suffer from two key limitations. First, their one-size-fits-all vector representations are…

Linguistic feature datasets such as URIEL+ are valuable for modelling cross-lingual relationships, but their high dimensionality and sparsity, especially for low-resource languages, limit the effectiveness of distance metrics. We propose a…

计算与语言 · 计算机科学 2025-09-25 York Hay Ng , Phuong Hanh Hoang , En-Shiun Annie Lee

In the pursuit of supporting more languages around the world, tools that characterize properties of languages play a key role in expanding the existing multilingual NLP research. In this study, we focus on a widely used typological…

计算与语言 · 计算机科学 2024-05-21 Hasti Toossi , Guo Qing Huai , Jinyu Liu , Eric Khiu , A. Seza Doğruöz , En-Shiun Annie Lee

Label projection is an effective technique for cross-lingual transfer, extending span-annotated datasets from a high-resource language to low-resource ones. Most approaches perform label projection as a separate step after machine…

计算与语言 · 计算机科学 2026-04-16 Thennal DK , Chris Biemann , Hans Ole Hatzel

India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognized as scheduled languages in the Indian Constitution. Despite…

This paper develops an approach to language identification in which the set of languages considered by the model depends on the geographic origin of the text in question. Given that many digital corpora can be geo-referenced at the country…

计算与语言 · 计算机科学 2024-03-18 Jonathan Dunn , Lane Edwards-Brown

Large pretrained multilingual models, trained on dozens of languages, have delivered promising results due to cross-lingual learning capabilities on variety of language tasks. Further adapting these models to specific languages, especially…

计算与语言 · 计算机科学 2022-11-24 Fahim Faisal , Antonios Anastasopoulos

Massively multilingual language models such as multilingual BERT offer state-of-the-art cross-lingual transfer performance on a range of NLP tasks. However, due to limited capacity and large differences in pretraining data sizes, there is a…

计算与语言 · 计算机科学 2021-09-13 Jonas Pfeiffer , Ivan Vulić , Iryna Gurevych , Sebastian Ruder

Although linguistic typology has a long history, computational approaches have only recently gained popularity. The use of distributed representations in computational linguistics has also become increasingly popular. A recent development…

计算与语言 · 计算机科学 2017-11-16 Johannes Bjerva , Isabelle Augenstein

Generative language modelling has surged in popularity with the emergence of services such as ChatGPT and Google Gemini. While these models have demonstrated transformative potential in productivity and communication, they overwhelmingly…

计算与语言 · 计算机科学 2025-07-09 Josh McGiff , Nikola S. Nikolov

Although multilingual LLMs have achieved remarkable performance across benchmarks, we find they continue to underperform on non-Latin script languages across contemporary LLM families. This discrepancy arises from the fact that LLMs are…

计算与语言 · 计算机科学 2025-06-27 Hoang H Nguyen , Khyati Mahajan , Vikas Yadav , Julian Salazar , Philip S. Yu , Masoud Hashemi , Rishabh Maheshwary

Large language models (LLMs) have achieved state-of-the-art performance in various software engineering tasks, including error detection, clone detection, and code translation, primarily leveraging high-resource programming languages like…

计算与语言 · 计算机科学 2025-06-11 Razan Baltaji , Saurabh Pujar , Louis Mandel , Martin Hirzel , Luca Buratti , Lav Varshney

Handwritten Text Recognition (HTR) under limited labeled data remains a challenging problem, particularly for Arabic-script languages. Although modern sequence-based recognizers perform well in high-resource settings, their accuracy…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Sana Al-azzawi , Elisa Barney , Marcus Liwicki

While generative multilingual models are rapidly being deployed, their safety and fairness evaluations are largely limited to resources collected in English. This is especially problematic for evaluations targeting inherently socio-cultural…

计算与语言 · 计算机科学 2024-03-12 Mukul Bhutani , Kevin Robinson , Vinodkumar Prabhakaran , Shachi Dave , Sunipa Dev

Script diversity presents a challenge to Multilingual Language Models (MLLM) by reducing lexical overlap among closely related languages. Therefore, transliterating closely related languages that use different writing scripts to a common…

计算与语言 · 计算机科学 2023-08-01 Ibraheem Muhammad Moosa , Mahmud Elahi Akhter , Ashfia Binte Habib

This paper describes a method to enrich lexical resources with content relating to linguistic diversity, based on knowledge from the field of lexical typology. We capture the phenomenon of diversity through the notions of lexical gap and…

Providing better language tools for low-resource and endangered languages is imperative for equitable growth. Recent progress with massively multilingual pretrained models has proven surprisingly effective at performing zero-shot transfer…

计算与语言 · 计算机科学 2022-11-10 Louis Clouâtre , Prasanna Parthasarathi , Amal Zouaq , Sarath Chandar

Cross-lingual representations have the potential to make NLP techniques available to the vast majority of languages in the world. However, they currently require large pretraining corpora or access to typologically similar languages. In…

计算与语言 · 计算机科学 2021-06-22 Wei Zhao , Steffen Eger , Johannes Bjerva , Isabelle Augenstein

The world's more than 7000 languages are written in at least 293 scripts. Due to various reasons, many closely related languages use different scripts, which poses a difficulty for multilingual pretrained language models (mPLMs) in learning…

计算与语言 · 计算机科学 2024-05-24 Yihong Liu , Chunlan Ma , Haotian Ye , Hinrich Schütze
‹ 上一页 1 2 3 10 下一页 ›