Modality Matching Matters: Calibrating Language Distances for Cross-Lingual Transfer in URIEL+
Abstract
Existing linguistic knowledge bases such as URIEL+ provide valuable geographic, genetic and typological distances for cross-lingual transfer but suffer from two key limitations. First, their one-size-fits-all vector representations are ill-suited to the diverse structures of linguistic data. Second, they lack a principled method for aggregating these signals into a single, comprehensive score. In this paper, we address these gaps by introducing a framework for type-matched language distances. We propose novel, structure-aware representations for each distance type: speaker-weighted distributions for geography, hyperbolic embeddings for genealogy, and a latent variables model for typology. We unify these signals into a robust, task-agnostic composite distance. Across multiple zero-shot transfer benchmarks, we demonstrate that our representations significantly improve transfer performance when the distance type is relevant to the task, while our composite distance yields gains in most tasks.
Cite
@article{arxiv.2510.19217,
title = {Modality Matching Matters: Calibrating Language Distances for Cross-Lingual Transfer in URIEL+},
author = {York Hay Ng and Aditya Khan and Xiang Lu and Matteo Salloum and Michael Zhou and Phuong H. Hoang and A. Seza Doğruöz and En-Shiun Annie Lee},
journal= {arXiv preprint arXiv:2510.19217},
year = {2026}
}
Comments
Accepted to EACL 2026 SRW