English

On the Difference of BERT-style and CLIP-style Text Encoders

Computation and Language 2023-06-07 v1

Abstract

Masked language modeling (MLM) has been one of the most popular pretraining recipes in natural language processing, e.g., BERT, one of the representative models. Recently, contrastive language-image pretraining (CLIP) has also attracted attention, especially its vision models that achieve excellent performance on a broad range of vision tasks. However, few studies are dedicated to studying the text encoders learned by CLIP. In this paper, we analyze the difference between BERT-style and CLIP-style text encoders from three experiments: (i) general text understanding, (ii) vision-centric text understanding, and (iii) text-to-image generation. Experimental analyses show that although CLIP-style text encoders underperform BERT-style ones for general text understanding tasks, they are equipped with a unique ability, i.e., synesthesia, for the cross-modal association, which is more similar to the senses of humans.

Keywords

Cite

@article{arxiv.2306.03678,
  title  = {On the Difference of BERT-style and CLIP-style Text Encoders},
  author = {Zhihong Chen and Guiming Hardy Chen and Shizhe Diao and Xiang Wan and Benyou Wang},
  journal= {arXiv preprint arXiv:2306.03678},
  year   = {2023}
}

Comments

Natural Language Processing. 10 pages, 1 figure. Findings of ACL-2023