中文

烹饪语境下的跨模态检索:学习语义文本-图像嵌入

计算与语言 2018-05-01 v1 计算机视觉与模式识别 信息检索

摘要

由于可用数据量巨大,加之能够分析这些数据的机器学习最新进展,设计支持烹饪活动的强大工具迅速普及。本文提出了一种跨模态检索模型,将视觉与文本数据(如菜肴图片及其食谱)对齐到共享表示空间中。我们描述了一种有效的学习方案,能够处理大规模问题,并在包含近 100 万张图片-食谱对的 Recipe1M 数据集上进行了验证。我们展示了我们的方法相对于先前 SOTA 模型的有效性,并给出了计算烹饪用例的定性结果。

关键词

引用

@article{arxiv.1804.11146,
  title  = {Cross-Modal Retrieval in the Cooking Context: Learning Semantic Text-Image Embeddings},
  author = {Micael Carvalho and Rémi Cadène and David Picard and Laure Soulier and Nicolas Thome and Matthieu Cord},
  journal= {arXiv preprint arXiv:1804.11146},
  year   = {2018}
}

备注

accepted at the 41st International ACM SIGIR Conference on Research and Development in Information Retrieval, 2018