English

Learning Multilingual Representation for Natural Language Understanding with Enhanced Cross-Lingual Supervision

Computation and Language 2021-06-10 v1

Abstract

Recently, pre-training multilingual language models has shown great potential in learning multilingual representation, a crucial topic of natural language processing. Prior works generally use a single mixed attention (MA) module, following TLM (Conneau and Lample, 2019), for attending to intra-lingual and cross-lingual contexts equivalently and simultaneously. In this paper, we propose a network named decomposed attention (DA) as a replacement of MA. The DA consists of an intra-lingual attention (IA) and a cross-lingual attention (CA), which model intralingual and cross-lingual supervisions respectively. In addition, we introduce a language-adaptive re-weighting strategy during training to further boost the model's performance. Experiments on various cross-lingual natural language understanding (NLU) tasks show that the proposed architecture and learning strategy significantly improve the model's cross-lingual transferability.

Keywords

Cite

@article{arxiv.2106.05166,
  title  = {Learning Multilingual Representation for Natural Language Understanding with Enhanced Cross-Lingual Supervision},
  author = {Yinpeng Guo and Liangyou Li and Xin Jiang and Qun Liu},
  journal= {arXiv preprint arXiv:2106.05166},
  year   = {2021}
}
R2 v1 2026-06-24T03:00:58.496Z