English

Order-Embeddings of Images and Language

Machine Learning 2016-03-02 v6 Computation and Language Computer Vision and Pattern Recognition

Abstract

Hypernymy, textual entailment, and image captioning can be seen as special cases of a single visual-semantic hierarchy over words, sentences, and images. In this paper we advocate for explicitly modeling the partial order structure of this hierarchy. Towards this goal, we introduce a general method for learning ordered representations, and show how it can be applied to a variety of tasks involving images and language. We show that the resulting representations improve performance over current approaches for hypernym prediction and image-caption retrieval.

Keywords

Cite

@article{arxiv.1511.06361,
  title  = {Order-Embeddings of Images and Language},
  author = {Ivan Vendrov and Ryan Kiros and Sanja Fidler and Raquel Urtasun},
  journal= {arXiv preprint arXiv:1511.06361},
  year   = {2016}
}

Comments

ICLR camera-ready version