English

Joint Learning of Distributed Representations for Images and Texts

Computer Vision and Pattern Recognition 2015-04-29 v2

Abstract

This technical report provides extra details of the deep multimodal similarity model (DMSM) which was proposed in (Fang et al. 2015, arXiv:1411.4952). The model is trained via maximizing global semantic similarity between images and their captions in natural language using the public Microsoft COCO database, which consists of a large set of images and their corresponding captions. The learned representations attempt to capture the combination of various visual concepts and cues.

Keywords

Cite

@article{arxiv.1504.03083,
  title  = {Joint Learning of Distributed Representations for Images and Texts},
  author = {Xiaodong He and Rupesh Srivastava and Jianfeng Gao and Li Deng},
  journal= {arXiv preprint arXiv:1504.03083},
  year   = {2015}
}

Comments

This is a previous tech report of a part of the work of arXiv:1411.4952. In order to avoid confusion, we'd like to withdraw this report from arXiv

R2 v1 2026-06-22T09:14:54.937Z