English

MiLiC-Eval: Benchmarking Multilingual LLMs for China's Minority Languages

Computation and Language 2025-06-03 v2

Abstract

Large language models (LLMs) excel in high-resource languages but struggle with low-resource languages (LRLs), particularly those spoken by minority communities in China, such as Tibetan, Uyghur, Kazakh, and Mongolian. To systematically track the progress in these languages, we introduce MiLiC-Eval, a benchmark designed for minority languages in China, featuring 24K instances across 9 tasks. MiLiC-Eval focuses on underrepresented writing systems. Its parallelism between tasks and languages can provide a faithful and fine-grained assessment of linguistic and problem-solving skills. Our evaluation reveals that open-source LLMs perform poorly on syntax-intensive tasks and multi-script languages. We further demonstrate how MiLiC-Eval can help advance LRL research in handling diverse writing systems and understanding the process of language adaptation.

Keywords

Cite

@article{arxiv.2503.01150,
  title  = {MiLiC-Eval: Benchmarking Multilingual LLMs for China's Minority Languages},
  author = {Chen Zhang and Mingxu Tao and Zhiyuan Liao and Yansong Feng},
  journal= {arXiv preprint arXiv:2503.01150},
  year   = {2025}
}

Comments

ACL 2025 (Findings) Code and data available at https://github.com/luciusssss/MiLiC-Eval

R2 v1 2026-06-28T22:04:02.759Z