中文

面向视觉与文本的复合混合表示

计算机视觉与模式识别 2022-06-15 v1 人工智能 机器学习

摘要

学习视觉与语言之间的公共表示空间可使深度网络将图像中的物体与相应语义关联起来。我们提出了一种模型,该模型学习共享的高斯混合表示,在无需显式位置监督的情况下将文本的组合性施加到视觉域。通过将空间变换器与表示学习方法相结合,我们学会将图像分割为分别编码的补丁,以可解释的方式关联视觉与文本表示。在 MNIST 和 CIFAR10 的变体上,我们的模型能够执行弱监督目标检测,并展示了其泛化到未见物体组合的能力。

关键词

引用

@article{arxiv.2206.06404,
  title  = {Compositional Mixture Representations for Vision and Text},
  author = {Stephan Alaniz and Marco Federici and Zeynep Akata},
  journal= {arXiv preprint arXiv:2206.06404},
  year   = {2022}
}

备注

Workshop on Learning with Limited Labelled Data for Image and Video Understanding (L3D-IVU), CVPR 2022