English

Count, Crop and Recognise: Fine-Grained Recognition in the Wild

Computer Vision and Pattern Recognition 2019-10-10 v2

Abstract

The goal of this paper is to label all the animal individuals present in every frame of a video. Unlike previous methods that have principally concentrated on labelling face tracks, we aim to label individuals even when their faces are not visible. We make the following contributions: (i) we introduce a 'Count, Crop and Recognise' (CCR) multistage recognition process for frame level labelling. The Count and Recognise stages involve specialised CNNs for the task, and we show that this simple staging gives a substantial boost in performance; (ii) we compare the recall using frame based labelling to both face and body track based labelling, and demonstrate the advantage of frame based with CCR for the specified goal; (iii) we introduce a new dataset for chimpanzee recognition in the wild; and (iv) we apply a high-granularity visualisation technique to further understand the learned CNN features for the recognition of chimpanzee individuals.

Keywords

Cite

@article{arxiv.1909.08950,
  title  = {Count, Crop and Recognise: Fine-Grained Recognition in the Wild},
  author = {Max Bain and Arsha Nagrani and Daniel Schofield and Andrew Zisserman},
  journal= {arXiv preprint arXiv:1909.08950},
  year   = {2019}
}
R2 v1 2026-06-23T11:20:10.859Z