English

UC2: Universal Cross-lingual Cross-modal Vision-and-Language Pre-training

Computer Vision and Pattern Recognition 2021-04-02 v1

Abstract

Vision-and-language pre-training has achieved impressive success in learning multimodal representations between vision and language. To generalize this success to non-English languages, we introduce UC2, the first machine translation-augmented framework for cross-lingual cross-modal representation learning. To tackle the scarcity problem of multilingual captions for image datasets, we first augment existing English-only datasets with other languages via machine translation (MT). Then we extend the standard Masked Language Modeling and Image-Text Matching training objectives to multilingual setting, where alignment between different languages is captured through shared visual context (i.e, using image as pivot). To facilitate the learning of a joint embedding space of images and all languages of interest, we further propose two novel pre-training tasks, namely Masked Region-to-Token Modeling (MRTM) and Visual Translation Language Modeling (VTLM), leveraging MT-enhanced translated data. Evaluation on multilingual image-text retrieval and multilingual visual question answering benchmarks demonstrates that our proposed framework achieves new state-of-the-art on diverse non-English benchmarks while maintaining comparable performance to monolingual pre-trained models on English tasks.

Keywords

Cite

@article{arxiv.2104.00332,
  title  = {UC2: Universal Cross-lingual Cross-modal Vision-and-Language Pre-training},
  author = {Mingyang Zhou and Luowei Zhou and Shuohang Wang and Yu Cheng and Linjie Li and Zhou Yu and Jingjing Liu},
  journal= {arXiv preprint arXiv:2104.00332},
  year   = {2021}
}
R2 v1 2026-06-24T00:45:55.630Z