English

Multi-scale Object-Aware Gaze Estimation via Geometric Reasoning

Computer Vision and Pattern Recognition 2026-06-28 v1

Abstract

Gaze target estimation aims to predict the semantic object an observer fixates upon within an image, a task deeply rooted in the object-oriented nature of human gaze. Observers tend to select a specific semantic entity as the attentional target, rather than responding randomly across arbitrary regions of the image. However, existing methods typically model this task as a direct mapping from global features to gaze heatmaps, essentially treating it as a pixel-level regression problem. This approach fails to explicitly represent the gazed object as a distinct entity, making it difficult to produce stable and semantically consistent predictions in complex scenes. To address this, we propose a two-stage gaze estimation framework guided by object semantics, reformulating gaze target estimation as a hierarchical reasoning process. Our method incorporates object-level representations during feature encoding to align image features with discrete semantic entities, then introduces multi-scale feature fusion and geometric constraints from head pose and gaze direction for fine-grained localization and object-level discrimination. Extensive experiments on GazeFollow, VideoAttentionTarget, ChildPlay, and GOO-Real demonstrate that our method achieves AUC of 0.961, 0.948, 0.987, and 0.977 respectively, delivering strong performance across all benchmarks while maintaining a compact parameter size of 7.1M.

Cite

@article{arxiv.2606.29334,
  title  = {Multi-scale Object-Aware Gaze Estimation via Geometric Reasoning},
  author = {Jiajie Mi and Xinyu Liu and Mengke Song and Chenglizhao Chen},
  journal= {arXiv preprint arXiv:2606.29334},
  year   = {2026}
}

Comments

Accepted by ECCV 2026