English

Learning to embed semantic similarity for joint image-text retrieval

Computer Vision and Pattern Recognition 2022-10-11 v1

Abstract

We present a deep learning approach for learning the joint semantic embeddings of images and captions in a Euclidean space, such that the semantic similarity is approximated by the L2 distances in the embedding space. For that, we introduce a metric learning scheme that utilizes multitask learning to learn the embedding of identical semantic concepts using a center loss. By introducing a differentiable quantization scheme into the end-to-end trainable network, we derive a semantic embedding of semantically similar concepts in Euclidean space. We also propose a novel metric learning formulation using an adaptive margin hinge loss, that is refined during the training phase. The proposed scheme was applied to the MS-COCO, Flicke30K and Flickr8K datasets, and was shown to compare favorably with contemporary state-of-the-art approaches.

Keywords

Cite

@article{arxiv.2210.03838,
  title  = {Learning to embed semantic similarity for joint image-text retrieval},
  author = {Noam Malali and Yosi Keller},
  journal= {arXiv preprint arXiv:2210.03838},
  year   = {2022}
}

Comments

in IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023

R2 v1 2026-06-28T03:02:29.780Z