English

Generating Diverse and Meaningful Captions

Computer Vision and Pattern Recognition 2018-12-20 v1 Computation and Language Machine Learning

Abstract

Image Captioning is a task that requires models to acquire a multi-modal understanding of the world and to express this understanding in natural language text. While the state-of-the-art for this task has rapidly improved in terms of n-gram metrics, these models tend to output the same generic captions for similar images. In this work, we address this limitation and train a model that generates more diverse and specific captions through an unsupervised training approach that incorporates a learning signal from an Image Retrieval model. We summarize previous results and improve the state-of-the-art on caption diversity and novelty. We make our source code publicly available online.

Keywords

Cite

@article{arxiv.1812.08126,
  title  = {Generating Diverse and Meaningful Captions},
  author = {Annika Lindh and Robert J. Ross and Abhijit Mahalunkar and Giancarlo Salton and John D. Kelleher},
  journal= {arXiv preprint arXiv:1812.08126},
  year   = {2018}
}

Comments

Accepted for presentation at The 27th International Conference on Artificial Neural Networks (ICANN 2018)

R2 v1 2026-06-23T06:48:15.046Z