English
Related papers

Related papers: Gaze-Driven Adaptive Interventions for Magazine-St…

200 papers

The emergence of advanced multimodal large language models (MLLMs) has significantly enhanced AI assistants' ability to process complex information across modalities. Recently, egocentric videos, by directly capturing user focus, actions,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Taiying Peng , Jiacheng Hua , Miao Liu , Feng Lu

We study the problem of inferring readers' identities and estimating their level of text comprehension from observations of their eye movements during reading. We develop a generative model of individual gaze patterns (scanpaths) that makes…

Machine Learning · Computer Science 2018-09-24 Silvia Makowski , Lena Jäger , Ahmed Abdelwahab , Niels Landwehr , Tobias Scheffer

Human intention is an internal, mental characterization for acquiring desired information. From interactive interfaces containing either textual or graphical information, intention to perceive desired information is subjective and strongly…

Human-Computer Interaction · Computer Science 2022-07-07 Shahed Anzarus Sabab , Mohammad Ridwan Kabir , Sayed Rizban Hussain , Hasan Mahmud , Md. Kamrul Hasan , Husne Ara Rubaiyeat

Interfaces for human oversight must effectively support users' situation awareness under time-critical conditions. We explore reinforcement learning (RL)-based UI adaptation to personalize alerting strategies that balance the benefits of…

Human-Computer Interaction · Computer Science 2026-02-10 Thorsten Klößner , João Belo , Zekun Wu , Jörg Hoffmann , Anna Maria Feit

This paper is interested in investigating whether human gaze signals can be leveraged to improve state-of-the-art search engine performance and how to incorporate this new input signal marked by human attention into existing neural…

Information Retrieval · Computer Science 2022-07-06 Sibo Dong , Justin Goldstein , Grace Hui Yang

A major challenge for physically unconstrained gaze estimation is acquiring training data with 3D gaze annotations for in-the-wild and outdoor scenarios. In contrast, videos of human interactions in unconstrained environments are abundantly…

Computer Vision and Pattern Recognition · Computer Science 2021-05-21 Rakshit Kothari , Shalini De Mello , Umar Iqbal , Wonmin Byeon , Seonwook Park , Jan Kautz

Empowering blind and low vision (BLV) users to explore visual media improves content comprehension, strengthens user agency, and fulfills diverse information needs. However, most existing tools separate exploration from the main narration,…

Human-Computer Interaction · Computer Science 2025-08-08 Shuchang Xu

Automatic eye gaze estimation is an important problem in vision based assistive technology with use cases in different emerging topics such as augmented reality, virtual reality and human-computer interaction. Over the past few years, there…

Computer Vision and Pattern Recognition · Computer Science 2022-08-08 Neeru Dubey , Shreya Ghosh , Abhinav Dhall

Visual navigation for autonomous agents is a core task in the fields of computer vision and robotics. Learning-based methods, such as deep reinforcement learning, have the potential to outperform the classical solutions developed for this…

Computer Vision and Pattern Recognition · Computer Science 2021-03-23 Zachary Seymour , Kowshik Thopalli , Niluthpol Mithun , Han-Pang Chiu , Supun Samarasekera , Rakesh Kumar

Modern information querying systems are progressively incorporating multimodal inputs like vision and audio. However, the integration of gaze -- a modality deeply linked to user intent and increasingly accessible via gaze-tracking wearables…

Human-Computer Interaction · Computer Science 2024-05-14 Zeyu Wang , Yuanchun Shi , Yuntao Wang , Yuchen Yao , Kun Yan , Yuhan Wang , Lei Ji , Xuhai Xu , Chun Yu

Nonverbal behaviors, particularly gaze direction, play a crucial role in enhancing effective communication in social interactions. As social robots increasingly participate in these interactions, they must adapt their gaze based on human…

Robotics · Computer Science 2026-02-13 Faezeh Vahedi , Morteza Memari , Ramtin Tabatabaei , Alireza Taheri

Eye-tracking offers rich insights into student cognition and engagement, but remains underutilized in classroom-facing educational technology due to challenges in data interpretation and accessibility. In this paper, we present the…

Recent advances in deep generative models demonstrate unprecedented zero-shot generalization capabilities, offering great potential for robot manipulation in unstructured environments. Given a partial observation of a scene, deep generative…

Robotics · Computer Science 2025-05-16 Le Shi , Yifei Shi , Xin Xu , Tenglong Liu , Junhua Xi , Chengyuan Chen

Artificial agents that support human group interactions hold great promise, especially in sensitive contexts such as well-being promotion and therapeutic interventions. However, current systems struggle to mediate group interactions…

Human-Computer Interaction · Computer Science 2026-03-17 Giulia Huang , Maristella Matera , Micol Spitale

In this paper, we address the intricate challenge of gaze vector prediction, a pivotal task with applications ranging from human-computer interaction to driver monitoring systems. Our innovative approach is designed for the demanding…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Abeer Banerjee , Naval K. Mehta , Shyam S. Prasad , Himanshu , Sumeet Saurav , Sanjay Singh

Predicting the target of visual search from eye fixation (gaze) data is a challenging problem with many applications in human-computer interaction. In contrast to previous work that has focused on individual instances as a search target, we…

Computer Vision and Pattern Recognition · Computer Science 2017-04-04 Hosnieh Sattar , Andreas Bulling , Mario Fritz

Visual Reinforcement Learning (RL) agents must learn to act based on high-dimensional image data where only a small fraction of the pixels is task-relevant. This forces agents to waste exploration and computational resources on irrelevant…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Andrew Lee , Ian Chuang , Dechen Gao , Kai Fukazawa , Iman Soltani

Current multimodal large language models (MLLMs) cannot effectively utilize eye-gaze information for video understanding, even when gaze cues are supplied via visual overlays or text descriptions. We introduce GazeQwen, a parameter…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Trong Thang Pham , Hien Nguyen , Ngan Le

Instructors often rely on visual actions such as pointing, marking, and sketching to convey information in educational presentation videos. These subtle visual cues often lack verbal descriptions, forcing low-vision (LV) learners to search…

Human-Computer Interaction · Computer Science 2025-08-06 Yotam Sechayk , Ariel Shamir , Amy Pavel , Takeo Igarashi

Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream tasks such as retrieval, classification, and segmentation. However, such representations tend to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Jona Ruthardt , Manu Gaur , Deva Ramanan , Makarand Tapaswi , Yuki M. Asano