English
Related papers

Related papers: Gaze-Regularized VLMs for Ego-Centric Behavior Und…

200 papers

An embodied AI assistant operating on egocentric video must integrate spatial cues across time - for instance, determining where an object A, glimpsed a few moments ago lies relative to an object B encountered later. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Sahithya Ravi , Gabriel Sarch , Vibhav Vineet , Andrew D. Wilson , Balasaravanan Thoravi Kumaravel

Gaze event detection is fundamental to vision science, human-computer interaction, and applied analytics. However, current workflows often require specialized programming knowledge and careful handling of heterogeneous raw data formats.…

Human-Computer Interaction · Computer Science 2026-04-16 Dongyang Guo , Yasmeen Abdrabou , Enkelejda Kasneci

Interpretable driver attention prediction is crucial for human-like autonomous driving. However, existing datasets provide only scene-level global gaze rather than fine-grained object-level annotations, inherently failing to support…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Zehong Ke , Yanbo Jiang , Jinhao Li , Zhiyuan Liu , Yiqian Tu , Qingwen Meng , Heye Huang , Jianqiang Wang

While exploring visual scenes, humans' scanpaths are driven by their underlying attention processes. Understanding visual scanpaths is essential for various applications. Traditional scanpath models predict the where and when of gaze shifts…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Xianyu Chen , Ming Jiang , Qi Zhao

Recent large-scale vision-language models (VLMs) have demonstrated remarkable capabilities in understanding and generating textual descriptions for visual content. However, these models lack an understanding of user-specific concepts. In…

Computer Vision and Pattern Recognition · Computer Science 2024-03-22 Yuval Alaluf , Elad Richardson , Sergey Tulyakov , Kfir Aberman , Daniel Cohen-Or

Gaze stabilization is critical for enabling fluid, accurate, and efficient interaction in immersive augmented reality (AR) environments, particularly during task-oriented visual behaviors. However, fixation sequences captured in active gaze…

Human-Computer Interaction · Computer Science 2025-10-03 Yaozheng Xia , Zaiping Zhu , Bo Pang , Shaorong Wang , Sheng Li

The visual focus of attention (VFOA) has been recognized as a prominent conversational cue. We are interested in estimating and tracking the VFOAs associated with multi-party social interactions. We note that in this type of situations the…

Computer Vision and Pattern Recognition · Computer Science 2018-12-21 Benoît Massé , Silèye Ba , Radu Horaud

The applications of Vision-Language Models (VLMs) in the field of Autonomous Driving (AD) have attracted widespread attention due to their outstanding performance and the ability to leverage Large Language Models (LLMs). By incorporating…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Xingcheng Zhou , Mingyu Liu , Ekim Yurtsever , Bare Luka Zagar , Walter Zimmer , Hu Cao , Alois C. Knoll

In this paper, we present a comprehensive and systematic analysis of vision-language models (VLMs) for disparate meme classification tasks. We introduced a novel approach that generates a VLM-based understanding of meme images and…

Computation and Language · Computer Science 2025-05-28 Deepesh Gavit , Debajyoti Mazumder , Samiran Das , Jasabanta Patro

In this paper a review is presented of the research on eye gaze estimation techniques and applications, that has progressed in diverse ways over the past two decades. Several generic eye gaze use-cases are identified: desktop, TV,…

Human-Computer Interaction · Computer Science 2017-08-08 Anuradha Kar , Peter Corcoran

Vision Language Models (VLMs) have achieved strong performance across diverse video understanding tasks. However, their viewpoint invariant training limits their ability to understand egocentric properties (e.g., human object interactions)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Dominick Reilly , Manish Kumar Govind , Le Xue , Srijan Das

In high-stakes domains, small task-specific vision models are crucial due to their low computational requirements and the availability of numerous methods to explain their results. However, these explanations often reveal that the models do…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Alexander Koebler , Lukas Kuhn , Ingo Thon , Florian Buettner

When intelligent agents learn visuomotor behaviors from human demonstrations, they may benefit from knowing where the human is allocating visual attention, which can be inferred from their gaze. A wealth of information regarding intelligent…

Computer Vision and Pattern Recognition · Computer Science 2018-06-12 Ruohan Zhang , Zhuode Liu , Luxin Zhang , Jake A. Whritner , Karl S. Muller , Mary M. Hayhoe , Dana H. Ballard

Rather than regressing gaze direction directly from images, we show that adding a 3D shape model can: i) improve gaze estimation accuracy, ii) perform well with lower resolution inputs and iii) provide a richer understanding of the…

Computer Vision and Pattern Recognition · Computer Science 2023-01-31 Hao Sun , Nick Pears

We introduce a novel self-improving framework that enhances Embodied Visual Tracking (EVT) with Vision-Language Models (VLMs) to address the limitations of current active visual tracking systems in recovering from tracking failure. Our…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Kui Wu , Shuhang Xu , Hao Chen , Churan Wang , Zhoujun Li , Yizhou Wang , Fangwei Zhong

A major challenge for physically unconstrained gaze estimation is acquiring training data with 3D gaze annotations for in-the-wild and outdoor scenarios. In contrast, videos of human interactions in unconstrained environments are abundantly…

Computer Vision and Pattern Recognition · Computer Science 2021-05-21 Rakshit Kothari , Shalini De Mello , Umar Iqbal , Wonmin Byeon , Seonwook Park , Jan Kautz

Gaze estimation involves predicting where the person is looking at within an image or video. Technically, the gaze information can be inferred from two different magnification levels: face orientation and eye orientation. The inference is…

Computer Vision and Pattern Recognition · Computer Science 2021-10-27 Ashesh , Chu-Song Chen , Hsuan-Tien Lin

Vision-Language Models (VLMs) offer the ability to generate high-level, interpretable descriptions of complex activities from images and videos, making them valuable for situational awareness (SA) applications. In such settings, the focus…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Pavana Pradeep , Krishna Kant , Suya Yu

In this work we employ multitask learning to capitalize on the structure that exists in related supervised tasks to train complex neural networks. It allows training a network for multiple objectives in parallel, in order to improve…

Computer Vision and Pattern Recognition · Computer Science 2019-09-17 Georgios Kapidis , Ronald Poppe , Elsbeth van Dam , Lucas Noldus , Remco Veltkamp

Data-driven saliency has recently gained a lot of attention thanks to the use of Convolutional Neural Networks for predicting gaze fixations. In this paper we go beyond standard approaches to saliency prediction, in which gaze maps are…

Computer Vision and Pattern Recognition · Computer Science 2018-07-10 Marcella Cornia , Lorenzo Baraldi , Giuseppe Serra , Rita Cucchiara