图像与文本分布式表示的联合学习
计算机视觉与模式识别
2015-04-29 v2
摘要
本技术报告提供了深度多模态相似度模型 (DMSM) 的额外细节,该模型由 (Fang et al. 2015, arXiv:1411.4952) 提出。该模型通过在公共 Microsoft COCO 数据库上最大化图像及其自然语言描述之间的全局语义相似性进行训练,该数据库包含大量图像及其对应的描述。学习到的表示试图捕捉各种视觉概念和线索的组合。
引用
@article{arxiv.1504.03083,
title = {Joint Learning of Distributed Representations for Images and Texts},
author = {Xiaodong He and Rupesh Srivastava and Jianfeng Gao and Li Deng},
journal= {arXiv preprint arXiv:1504.03083},
year = {2015}
}
备注
This is a previous tech report of a part of the work of arXiv:1411.4952. In order to avoid confusion, we'd like to withdraw this report from arXiv