English

Contrastive Language-Image Pre-training for the Italian Language

Computation and Language 2021-08-20 v1 Computer Vision and Pattern Recognition

Abstract

CLIP (Contrastive Language-Image Pre-training) is a very recent multi-modal model that jointly learns representations of images and texts. The model is trained on a massive amount of English data and shows impressive performance on zero-shot classification tasks. Training the same model on a different language is not trivial, since data in other languages might be not enough and the model needs high-quality translations of the texts to guarantee a good performance. In this paper, we present the first CLIP model for the Italian Language (CLIP-Italian), trained on more than 1.4 million image-text pairs. Results show that CLIP-Italian outperforms the multilingual CLIP model on the tasks of image retrieval and zero-shot classification.

Keywords

Cite

@article{arxiv.2108.08688,
  title  = {Contrastive Language-Image Pre-training for the Italian Language},
  author = {Federico Bianchi and Giuseppe Attanasio and Raphael Pisoni and Silvia Terragni and Gabriele Sarti and Sri Lakshmi},
  journal= {arXiv preprint arXiv:2108.08688},
  year   = {2021}
}
R2 v1 2026-06-24T05:15:11.859Z