English

Common Sense Beyond English: Evaluating and Improving Multilingual Language Models for Commonsense Reasoning

Computation and Language 2021-06-15 v1 Artificial Intelligence

Abstract

Commonsense reasoning research has so far been limited to English. We aim to evaluate and improve popular multilingual language models (ML-LMs) to help advance commonsense reasoning (CSR) beyond English. We collect the Mickey Corpus, consisting of 561k sentences in 11 different languages, which can be used for analyzing and improving ML-LMs. We propose Mickey Probe, a language-agnostic probing task for fairly evaluating the common sense of popular ML-LMs across different languages. In addition, we also create two new datasets, X-CSQA and X-CODAH, by translating their English versions to 15 other languages, so that we can evaluate popular ML-LMs for cross-lingual commonsense reasoning. To improve the performance beyond English, we propose a simple yet effective method -- multilingual contrastive pre-training (MCP). It significantly enhances sentence representations, yielding a large performance gain on both benchmarks.

Keywords

Cite

@article{arxiv.2106.06937,
  title  = {Common Sense Beyond English: Evaluating and Improving Multilingual Language Models for Commonsense Reasoning},
  author = {Bill Yuchen Lin and Seyeon Lee and Xiaoyang Qiao and Xiang Ren},
  journal= {arXiv preprint arXiv:2106.06937},
  year   = {2021}
}

Comments

Accepted to ACL-IJCNLP 2021 (long paper at main conference). Project website: https://inklab.usc.edu/XCSR/

R2 v1 2026-06-24T03:08:30.411Z