English

OpenFashionCLIP: Vision-and-Language Contrastive Learning with Open-Source Fashion Data

Computer Vision and Pattern Recognition 2023-09-12 v1

Abstract

The inexorable growth of online shopping and e-commerce demands scalable and robust machine learning-based solutions to accommodate customer requirements. In the context of automatic tagging classification and multimodal retrieval, prior works either defined a low generalizable supervised learning approach or more reusable CLIP-based techniques while, however, training on closed source data. In this work, we propose OpenFashionCLIP, a vision-and-language contrastive learning method that only adopts open-source fashion data stemming from diverse domains, and characterized by varying degrees of specificity. Our approach is extensively validated across several tasks and benchmarks, and experimental results highlight a significant out-of-domain generalization capability and consistent improvements over state-of-the-art methods both in terms of accuracy and recall. Source code and trained models are publicly available at: https://github.com/aimagelab/open-fashion-clip.

Keywords

Cite

@article{arxiv.2309.05551,
  title  = {OpenFashionCLIP: Vision-and-Language Contrastive Learning with Open-Source Fashion Data},
  author = {Giuseppe Cartella and Alberto Baldrati and Davide Morelli and Marcella Cornia and Marco Bertini and Rita Cucchiara},
  journal= {arXiv preprint arXiv:2309.05551},
  year   = {2023}
}

Comments

International Conference on Image Analysis and Processing (ICIAP) 2023