English

Harnessing Linguistic Dissimilarity for Language Generalization on Unseen Low-Resource Varieties

Computation and Language 2026-05-07 v1 Artificial Intelligence

Abstract

Low-resource language varieties used by specific groups remain neglected in the development of Multilingual Language Models. A great deal of cross-lingual research focuses on inter-lingual language transfer which strives to align allied varieties and minimize differences between them. However, for low-resource varieties, linguistic dissimilarity is also an important cue allowing generalization to unseen varieties. Unlike prior approaches, we propose a two-stage Language Generalization framework that focuses on capturing variety-specific cues while also exploiting rich overlap offered by high-resource source variety. First, we propose TOPPing, a source-selection method specifically designed for low-resource varieties. Second, we suggest a lightweight VACAI-Bowl architecture that learns variety-specific attributes with one branch while a parallel branch captures variety-invariant attributes using adversarial training. We evaluate our framework on structural prediction tasks, which are among the few tasks available, as proxy for performance on other downstream tasks. Using VACAI-Bowl with TOPPing yields an average 54.62% improvement in the dependency parsing task, which serves as a proxy for performance on other downstream tasks across 10 low-resource varieties.

Keywords

Cite

@article{arxiv.2605.04500,
  title  = {Harnessing Linguistic Dissimilarity for Language Generalization on Unseen Low-Resource Varieties},
  author = {Jinju Kim and Haeji Jung and Youjeong Roh and Jong Hwan Ko and David R. Mortensen},
  journal= {arXiv preprint arXiv:2605.04500},
  year   = {2026}
}

Comments

Accepted to CoNLL 2026

R2 v1 2026-07-01T12:52:09.784Z