English

A fast compression-based similarity measure with applications to content-based image retrieval

Machine Learning 2012-10-03 v1 Information Retrieval Machine Learning

Abstract

Compression-based similarity measures are effectively employed in applications on diverse data types with a basically parameter-free approach. Nevertheless, there are problems in applying these techniques to medium-to-large datasets which have been seldom addressed. This paper proposes a similarity measure based on compression with dictionaries, the Fast Compression Distance (FCD), which reduces the complexity of these methods, without degradations in performance. On its basis a content-based color image retrieval system is defined, which can be compared to state-of-the-art methods based on invariant color features. Through the FCD a better understanding of compression-based techniques is achieved, by performing experiments on datasets which are larger than the ones analyzed so far in literature.

Keywords

Cite

@article{arxiv.1210.0758,
  title  = {A fast compression-based similarity measure with applications to content-based image retrieval},
  author = {Daniele Cerra and Mihai Datcu},
  journal= {arXiv preprint arXiv:1210.0758},
  year   = {2012}
}

Comments

Pre-print

R2 v1 2026-06-21T22:14:39.858Z