English

Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation

Computer Vision and Pattern Recognition 2024-01-15 v1 Artificial Intelligence

Abstract

The process of transforming input images into corresponding textual explanations stands as a crucial and complex endeavor within the domains of computer vision and natural language processing. In this paper, we propose an innovative ensemble approach that harnesses the capabilities of Contrastive Language-Image Pretraining models.

Keywords

Cite

@article{arxiv.2401.06167,
  title  = {Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation},
  author = {Chang Che and Qunwei Lin and Xinyu Zhao and Jiaxin Huang and Liqiang Yu},
  journal= {arXiv preprint arXiv:2401.06167},
  year   = {2024}
}
R2 v1 2026-06-28T14:14:38.658Z