English

A Lightweight Feature Fusion Architecture For Resource-Constrained Crowd Counting

Computer Vision and Pattern Recognition 2024-01-12 v1

Abstract

Crowd counting finds direct applications in real-world situations, making computational efficiency and performance crucial. However, most of the previous methods rely on a heavy backbone and a complex downstream architecture that restricts the deployment. To address this challenge and enhance the versatility of crowd-counting models, we introduce two lightweight models. These models maintain the same downstream architecture while incorporating two distinct backbones: MobileNet and MobileViT. We leverage Adjacent Feature Fusion to extract diverse scale features from a Pre-Trained Model (PTM) and subsequently combine these features seamlessly. This approach empowers our models to achieve improved performance while maintaining a compact and efficient design. With the comparison of our proposed models with previously available state-of-the-art (SOTA) methods on ShanghaiTech-A ShanghaiTech-B and UCF-CC-50 dataset, it achieves comparable results while being the most computationally efficient model. Finally, we present a comparative study, an extensive ablation study, along with pruning to show the effectiveness of our models.

Keywords

Cite

@article{arxiv.2401.05968,
  title  = {A Lightweight Feature Fusion Architecture For Resource-Constrained Crowd Counting},
  author = {Yashwardhan Chaudhuri and Ankit Kumar and Orchid Chetia Phukan and Arun Balaji Buduru},
  journal= {arXiv preprint arXiv:2401.05968},
  year   = {2024}
}
R2 v1 2026-06-28T14:14:21.388Z