English

Alleviating Noisy Data in Image Captioning with Cooperative Distillation

Computer Vision and Pattern Recognition 2020-12-23 v1 Machine Learning

Abstract

Image captioning systems have made substantial progress, largely due to the availability of curated datasets like Microsoft COCO or Vizwiz that have accurate descriptions of their corresponding images. Unfortunately, scarce availability of such cleanly labeled data results in trained algorithms producing captions that can be terse and idiosyncratically specific to details in the image. We propose a new technique, cooperative distillation that combines clean curated datasets with the web-scale automatically extracted captions of the Google Conceptual Captions dataset (GCC), which can have poor descriptions of images, but is abundant in size and therefore provides a rich vocabulary resulting in more expressive captions.

Keywords

Cite

@article{arxiv.2012.11691,
  title  = {Alleviating Noisy Data in Image Captioning with Cooperative Distillation},
  author = {Pierre Dognin and Igor Melnyk and Youssef Mroueh and Inkit Padhi and Mattia Rigotti and Jarret Ross and Yair Schiff},
  journal= {arXiv preprint arXiv:2012.11691},
  year   = {2020}
}

Comments

CVPR 2020 VizWiz Challenge

R2 v1 2026-06-23T21:10:11.533Z