English

What if I ask in \textit{alia lingua}? Measuring Functional Similarity Across Languages

Computation and Language 2025-10-03 v2 Machine Learning

Abstract

How similar are model outputs across languages? In this work, we study this question using a recently proposed model similarity metric κp\kappa_p applied to 20 languages and 47 subjects in GlobalMMLU. Our analysis reveals that a model's responses become increasingly consistent across languages as its size and capability grow. Interestingly, models exhibit greater cross-lingual consistency within themselves than agreement with other models prompted in the same language. These results highlight not only the value of κp\kappa_p as a practical tool for evaluating multilingual reliability, but also its potential to guide the development of more consistent multilingual systems.

Keywords

Cite

@article{arxiv.2509.04032,
  title  = {What if I ask in \textit{alia lingua}? Measuring Functional Similarity Across Languages},
  author = {Debangan Mishra and Arihant Rastogi and Agyeya Negi and Shashwat Goel and Ponnurangam Kumaraguru},
  journal= {arXiv preprint arXiv:2509.04032},
  year   = {2025}
}

Comments

Accepted into Multilingual Representation Learning (MRL) Workshop at EMNLP 2025