烹饪语境下的跨模态检索:学习语义文本-图像嵌入
计算与语言
2018-05-01 v1 计算机视觉与模式识别
信息检索
摘要
由于可用数据量巨大,加之能够分析这些数据的机器学习最新进展,设计支持烹饪活动的强大工具迅速普及。本文提出了一种跨模态检索模型,将视觉与文本数据(如菜肴图片及其食谱)对齐到共享表示空间中。我们描述了一种有效的学习方案,能够处理大规模问题,并在包含近 100 万张图片-食谱对的 Recipe1M 数据集上进行了验证。我们展示了我们的方法相对于先前 SOTA 模型的有效性,并给出了计算烹饪用例的定性结果。
关键词
引用
@article{arxiv.1804.11146,
title = {Cross-Modal Retrieval in the Cooking Context: Learning Semantic Text-Image Embeddings},
author = {Micael Carvalho and Rémi Cadène and David Picard and Laure Soulier and Nicolas Thome and Matthieu Cord},
journal= {arXiv preprint arXiv:1804.11146},
year = {2018}
}
备注
accepted at the 41st International ACM SIGIR Conference on Research and Development in Information Retrieval, 2018