English

Few-Shot Adaptation for Multimedia Semantic Indexing

Multimedia 2018-07-20 v1 Computer Vision and Pattern Recognition

Abstract

We propose a few-shot adaptation framework, which bridges zero-shot learning and supervised many-shot learning, for semantic indexing of image and video data. Few-shot adaptation provides robust parameter estimation with few training examples, by optimizing the parameters of zero-shot learning and supervised many-shot learning simultaneously. In this method, first we build a zero-shot detector, and then update it by using the few examples. Our experiments show the effectiveness of the proposed framework on three datasets: TRECVID Semantic Indexing 2010, 2014, and ImageNET. On the ImageNET dataset, we show that our method outperforms recent few-shot learning methods. On the TRECVID 2014 dataset, we achieve 15.19% and 35.98% in Mean Average Precision under the zero-shot condition and the supervised condition, respectively. To the best of our knowledge, these are the best results on this dataset.

Keywords

Cite

@article{arxiv.1807.07203,
  title  = {Few-Shot Adaptation for Multimedia Semantic Indexing},
  author = {Nakamasa Inoue and Koichi Shinoda},
  journal= {arXiv preprint arXiv:1807.07203},
  year   = {2018}
}
R2 v1 2026-06-23T03:06:40.398Z