English

XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation

Computation and Language 2021-10-08 v2 Artificial Intelligence

Abstract

Machine learning has brought striking advances in multilingual natural language processing capabilities over the past year. For example, the latest techniques have improved the state-of-the-art performance on the XTREME multilingual benchmark by more than 13 points. While a sizeable gap to human-level performance remains, improvements have been easier to achieve in some tasks than in others. This paper analyzes the current state of cross-lingual transfer learning and summarizes some lessons learned. In order to catalyze meaningful progress, we extend XTREME to XTREME-R, which consists of an improved set of ten natural language understanding tasks, including challenging language-agnostic retrieval tasks, and covers 50 typologically diverse languages. In addition, we provide a massively multilingual diagnostic suite (MultiCheckList) and fine-grained multi-dataset evaluation capabilities through an interactive public leaderboard to gain a better understanding of such models. The leaderboard and code for XTREME-R will be made available at https://sites.research.google/xtreme and https://github.com/google-research/xtreme respectively.

Keywords

Cite

@article{arxiv.2104.07412,
  title  = {XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation},
  author = {Sebastian Ruder and Noah Constant and Jan Botha and Aditya Siddhant and Orhan Firat and Jinlan Fu and Pengfei Liu and Junjie Hu and Dan Garrette and Graham Neubig and Melvin Johnson},
  journal= {arXiv preprint arXiv:2104.07412},
  year   = {2021}
}

Comments

EMNLP 2021 camera-ready

R2 v1 2026-06-24T01:11:52.033Z