CurlingNet:面向 Fashion IQ 数据的图像与文本间组合学习
计算机视觉与模式识别
2020-03-31 v2
摘要
我们提出一种名为 CurlingNet 的方法,可度量图像-文本嵌入组合之间的语义距离。为针对时尚领域数据学习有效的图像-文本组合,我们的模型提出如下两个关键组件。第一,Delivery 在嵌入空间中对源图像进行过渡。第二,Sweeping 在嵌入空间中强调时尚图像中与查询相关的组件。我们利用通道级门控机制使之成为可能。我们的单一模型优于以往 state-of-the-art(最优)的图像-文本组合模型,包括 TIRG 与 FiLM。我们参与了 ICCV 2019 的首届 fashion-IQ 挑战赛,我们模型的集成取得了最佳表现之一。
引用
@article{arxiv.2003.12299,
title = {CurlingNet: Compositional Learning between Images and Text for Fashion IQ Data},
author = {Youngjae Yu and Seunghwan Lee and Yuncheol Choi and Gunhee Kim},
journal= {arXiv preprint arXiv:2003.12299},
year = {2020}
}
备注
4 pages, 4 figures, ICCV 2019 Linguistics Meets image and video retrieval workshop, Fashion IQ challenge