English

Localized Vision-Language Matching for Open-vocabulary Object Detection

Computer Vision and Pattern Recognition 2022-07-29 v2 Machine Learning

Abstract

In this work, we propose an open-vocabulary object detection method that, based on image-caption pairs, learns to detect novel object classes along with a given set of known classes. It is a two-stage training approach that first uses a location-guided image-caption matching technique to learn class labels for both novel and known classes in a weakly-supervised manner and second specializes the model for the object detection task using known class annotations. We show that a simple language model fits better than a large contextualized language model for detecting novel objects. Moreover, we introduce a consistency-regularization technique to better exploit image-caption pair information. Our method compares favorably to existing open-vocabulary detection approaches while being data-efficient. Source code is available at https://github.com/lmb-freiburg/locov .

Keywords

Cite

@article{arxiv.2205.06160,
  title  = {Localized Vision-Language Matching for Open-vocabulary Object Detection},
  author = {Maria A. Bravo and Sudhanshu Mittal and Thomas Brox},
  journal= {arXiv preprint arXiv:2205.06160},
  year   = {2022}
}

Comments

Accepted at DAGM German Conference on Pattern Recognition (GCPR 2022)

R2 v1 2026-06-24T11:15:38.070Z