English

MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation

Computation and Language 2025-06-04 v2 Artificial Intelligence

Abstract

With the rapid adoption of large language models (LLMs) in natural language processing, the ability to follow instructions has emerged as a key metric for evaluating their practical utility. However, existing evaluation methods often focus on single-language scenarios, overlooking the challenges and differences present in multilingual and cross-lingual contexts. To address this gap, we introduce MaXIFE: a comprehensive evaluation benchmark designed to assess instruction-following capabilities across 23 different languages with 1667 verifiable instruction tasks. MaXIFE integrates both Rule-Based Evaluation and Model-Based Evaluation, ensuring a balance of efficiency and accuracy. We applied MaXIFE to evaluate several leading commercial LLMs, establishing baseline results for future comparisons. By providing a standardized tool for multilingual instruction-following evaluation, MaXIFE aims to advance research and development in natural language processing.

Keywords

Cite

@article{arxiv.2506.01776,
  title  = {MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation},
  author = {Yile Liu and Ziwei Ma and Xiu Jiang and Jinglu Hu and Jing Chang and Liang Li},
  journal= {arXiv preprint arXiv:2506.01776},
  year   = {2025}
}

Comments

ACL 2025 Main Conference