English

Unified Attention Modeling for Efficient Free-Viewing and Visual Search via Shared Representations

Computer Vision and Pattern Recognition 2025-06-04 v1 Artificial Intelligence

Abstract

Computational human attention modeling in free-viewing and task-specific settings is often studied separately, with limited exploration of whether a common representation exists between them. This work investigates this question and proposes a neural network architecture that builds upon the Human Attention transformer (HAT) to test the hypothesis. Our results demonstrate that free-viewing and visual search can efficiently share a common representation, allowing a model trained in free-viewing attention to transfer its knowledge to task-driven visual search with a performance drop of only 3.86% in the predicted fixation scanpaths, measured by the semantic sequence score (SemSS) metric which reflects the similarity between predicted and human scanpaths. This transfer reduces computational costs by 92.29% in terms of GFLOPs and 31.23% in terms of trainable parameters.

Keywords

Cite

@article{arxiv.2506.02764,
  title  = {Unified Attention Modeling for Efficient Free-Viewing and Visual Search via Shared Representations},
  author = {Fatma Youssef Mohammed and Kostas Alexis},
  journal= {arXiv preprint arXiv:2506.02764},
  year   = {2025}
}

Comments

Accepted to the 2025 IEEE International Conference on Development and Learning (ICDL)

R2 v1 2026-07-01T02:56:42.908Z