English

Enhancing Vision-Language Pre-training with Rich Supervisions

Computer Vision and Pattern Recognition 2025-03-14 v2

Abstract

We propose Strongly Supervised pre-training with ScreenShots (S4) - a novel pre-training paradigm for Vision-Language Models using data from large-scale web screenshot rendering. Using web screenshots unlocks a treasure trove of visual and textual cues that are not present in using image-text pairs. In S4, we leverage the inherent tree-structured hierarchy of HTML elements and the spatial localization to carefully design 10 pre-training tasks with large scale annotated data. These tasks resemble downstream tasks across different domains and the annotations are cheap to obtain. We demonstrate that, compared to current screenshot pre-training objectives, our innovative pre-training method significantly enhances performance of image-to-text model in nine varied and popular downstream tasks - up to 76.1% improvements on Table Detection, and at least 1% on Widget Captioning.

Keywords

Cite

@article{arxiv.2403.03346,
  title  = {Enhancing Vision-Language Pre-training with Rich Supervisions},
  author = {Yuan Gao and Kunyu Shi and Pengkai Zhu and Edouard Belval and Oren Nuriel and Srikar Appalaraju and Shabnam Ghadar and Vijay Mahadevan and Zhuowen Tu and Stefano Soatto},
  journal= {arXiv preprint arXiv:2403.03346},
  year   = {2025}
}

Comments

Accepted to CVPR 2024

R2 v1 2026-06-28T15:10:25.166Z