English

MCSE: Multimodal Contrastive Learning of Sentence Embeddings

Computation and Language 2022-04-26 v1

Abstract

Learning semantically meaningful sentence embeddings is an open problem in natural language processing. In this work, we propose a sentence embedding learning approach that exploits both visual and textual information via a multimodal contrastive objective. Through experiments on a variety of semantic textual similarity tasks, we demonstrate that our approach consistently improves the performance across various datasets and pre-trained encoders. In particular, combining a small amount of multimodal data with a large text-only corpus, we improve the state-of-the-art average Spearman's correlation by 1.7%. By analyzing the properties of the textual embedding space, we show that our model excels in aligning semantically similar sentences, providing an explanation for its improved performance.

Keywords

Cite

@article{arxiv.2204.10931,
  title  = {MCSE: Multimodal Contrastive Learning of Sentence Embeddings},
  author = {Miaoran Zhang and Marius Mosbach and David Ifeoluwa Adelani and Michael A. Hedderich and Dietrich Klakow},
  journal= {arXiv preprint arXiv:2204.10931},
  year   = {2022}
}

Comments

Accepted by NAACL 2022 main conference (short paper), 11 pages