English

AA: A Multi-view Multimodal Dataset for Screen-based Gaze Estimation

Computer Vision and Pattern Recognition 2026-06-30 v1 Human-Computer Interaction

Abstract

We present AA, a multi-view multimodal dataset for screen-based gaze estimation. The dataset captures synchronized facial observations from eight fixed screen-mounted cameras and two additional side-view cameras, paired with precise screen-space gaze targets collected under controlled fixation conditions. Each sample contains multi-view face observations together with structured facial region crops, enabling multimodal learning from both global and local visual cues. Unlike existing single-view gaze datasets, AA provides multi-view coverage from both screen-mounted and side-mounted perspectives, enabling more robust modeling under viewpoint variation and occlusion. The dataset includes subject-independent evaluation splits and a standardized data processing pipeline to support reproducible research in gaze estimation.

Cite

@article{arxiv.2606.31211,
  title  = {AA: A Multi-view Multimodal Dataset for Screen-based Gaze Estimation},
  author = {Chang Liu and Jiaqi Liu and Zhoutong Ye and Xinjie Shen and Chun Yu and Yuanchun Shi},
  journal= {arXiv preprint arXiv:2606.31211},
  year   = {2026}
}