English

SynthRef: Generation of Synthetic Referring Expressions for Object Segmentation

Computer Vision and Pattern Recognition 2021-06-10 v2 Computation and Language Multimedia

Abstract

Recent advances in deep learning have brought significant progress in visual grounding tasks such as language-guided video object segmentation. However, collecting large datasets for these tasks is expensive in terms of annotation time, which represents a bottleneck. To this end, we propose a novel method, namely SynthRef, for generating synthetic referring expressions for target objects in an image (or video frame), and we also present and disseminate the first large-scale dataset with synthetic referring expressions for video object segmentation. Our experiments demonstrate that by training with our synthetic referring expressions one can improve the ability of a model to generalize across different datasets, without any additional annotation cost. Moreover, our formulation allows its application to any object detection or segmentation dataset.

Keywords

Cite

@article{arxiv.2106.04403,
  title  = {SynthRef: Generation of Synthetic Referring Expressions for Object Segmentation},
  author = {Ioannis Kazakos and Carles Ventura and Miriam Bellver and Carina Silberer and Xavier Giro-i-Nieto},
  journal= {arXiv preprint arXiv:2106.04403},
  year   = {2021}
}

Comments

Accepted as poster at the NAACL 2021 Visually Grounded Interaction and Language (ViGIL) Workshop. 4 pages. Project website: https://imatge-upc.github.io/synthref/

R2 v1 2026-06-24T02:57:46.498Z