English

Seeing through bag-of-visual-word glasses: towards understanding quantization effects in feature extraction methods

Computer Vision and Pattern Recognition 2014-08-21 v1

Abstract

Vector-quantized local features frequently used in bag-of-visual-words approaches are the backbone of popular visual recognition systems due to both their simplicity and their performance. Despite their success, bag-of-words-histograms basically contain low-level image statistics (e.g., number of edges of different orientations). The question remains how much visual information is "lost in quantization" when mapping visual features to code words? To answer this question, we present an in-depth analysis of the effect of local feature quantization on human recognition performance. Our analysis is based on recovering the visual information by inverting quantized local features and presenting these visualizations with different codebook sizes to human observers. Although feature inversion techniques are around for quite a while, to the best of our knowledge, our technique is the first visualizing especially the effect of feature quantization. Thereby, we are now able to compare single steps in common image classification pipelines to human counterparts.

Keywords

Cite

@article{arxiv.1408.4692,
  title  = {Seeing through bag-of-visual-word glasses: towards understanding quantization effects in feature extraction methods},
  author = {Alexander Freytag and Johannes Rühle and Paul Bodesheim and Erik Rodner and Joachim Denzler},
  journal= {arXiv preprint arXiv:1408.4692},
  year   = {2014}
}

Comments

An abstract version of this paper was accepted for the ICPR FEAST Workshop

R2 v1 2026-06-22T05:34:52.142Z