English
Related papers

Related papers: GazeSummary: Exploring Gaze as an Implicit Prompt …

200 papers

Current LLM assistants are powerful at answering questions, but they have limited access to the behavioral context that reveals when and where a user is struggling. We present a gaze-grounded multimodal LLM assistant that uses egocentric…

Human-Computer Interaction · Computer Science 2026-04-10 Valdemar Danry , Javier Hernandez , Andrew Wilson , Pattie Maes , Judith Amores

A promising effective human-robot interaction in assistive robotic systems is gaze-based control. However, current gaze-based assistive systems mainly help users with basic grasping actions, offering limited support. Moreover, the…

Robotics · Computer Science 2025-08-20 Zejia Zhang , Bo Yang , Xinxing Chen , Weizhuang Shi , Haoyuan Wang , Wei Luo , Jian Huang

Eye gaze offers valuable cues about attention, short-term intent, and future actions, making it a powerful signal for modeling egocentric behavior. In this work, we propose a gaze-regularized framework that enhances VLMs for two key…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Anupam Pani , Yanchao Yang

Smart glasses with AI assistants are increasingly used in daily life. However, current systems lack awareness of the user's internal cognitive state, leaving them unable to proactively anticipate users' needs without access to cognitive…

The ability to read, understand and find important information from written text is a critical skill in our daily lives for our independence, comfort and safety. However, a significant part of our society is affected by partial vision…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Wiktor Mucha , Florin Cuconasu , Naome A. Etori , Valia Kalokyri , Giovanni Trappolini

The emergence of advanced multimodal large language models (MLLMs) has significantly enhanced AI assistants' ability to process complex information across modalities. Recently, egocentric videos, by directly capturing user focus, actions,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Taiying Peng , Jiacheng Hua , Miao Liu , Feng Lu

Large Language Models (LLMs) are advancing into Multimodal LLMs (MLLMs), capable of processing image, audio, and video as well as text. Combining first-person video, MLLMs show promising potential for understanding human activities through…

Human-Computer Interaction · Computer Science 2025-04-09 Jun Rekimoto

Multimodal large language models (LMMs) excel in world knowledge and problem-solving abilities. Through the use of a world-facing camera and contextual AI, emerging smart accessories aim to provide a seamless interface between humans and…

Human-Computer Interaction · Computer Science 2024-02-01 Robert Konrad , Nitish Padmanaban , J. Gabriel Buckmaster , Kevin C. Boyle , Gordon Wetzstein

Eye gaze, encompassing fixations and saccades, provides critical insights into human intentions and future actions. This study introduces a gaze-regularized framework that enhances Vision Language Models (VLMs) for egocentric behavior…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Anupam Pani , Yanchao Yang

While being able to read with screen magnifiers, low vision people have slow and unpleasant reading experiences. Eye tracking has the potential to improve their experience by recognizing fine-grained gaze behaviors and providing more…

Human-Computer Interaction · Computer Science 2023-03-30 Ru Wang , Linxiu Zeng , Xinyong Zhang , Sanbrita Mondal , Yuhang Zhao

Large Language Models (LLMs) have substantially improved the conversational capabilities of social robots. Nevertheless, for an intuitive and fluent human-robot interaction, robots should be able to ground the conversation by relating…

Human-Computer Interaction · Computer Science 2026-04-09 Elisabeth Menendez , Michael Gienger , Santiago Martínez , Carlos Balaguer , Anna Belardinelli

Vision-language models (VLMs) have rapidly evolved into general-purpose multimodal reasoners with strong zero-shot generalization. In this context, VLMs could greatly benefit the analysis of human gaze and attention, a central task in human…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Hengfei Wang , Anshul Gupta , Pierre Vuillecard , Jean-Marc Odobez

Reading is a challenging task for low vision people. While conventional low vision aids (e.g., magnification) offer certain support, they cannot fully address the difficulties faced by low vision users, such as locating the next line and…

Human-Computer Interaction · Computer Science 2024-02-26 Ru Wang , Zach Potter , Yun Ho , Daniel Killough , Linxiu Zeng , Sanbrita Mondal , Yuhang Zhao

The rapid development of Large Language Models (LLMs) creates an exciting potential for flexible, general knowledge-driven Human-Robot Interaction (HRI) systems for assistive robots. Existing HRI systems demonstrate great progress in…

Robotics · Computer Science 2025-07-22 Jens V. Rüppel , Andrey Rudenko , Tim Schreiter , Martin Magnusson , Achim J. Lilienthal

Recent advancements in eye tracking technology are driving the adoption of gaze-assisted interaction as a rich and accessible human-computer interaction paradigm. Gaze-assisted interaction serves as a contextual, non-invasive, and explicit…

Human-Computer Interaction · Computer Science 2018-03-14 Vijay Rajanna , Tracy Hammond

Comprehensively interpreting human behavior is a core challenge in human-aware artificial intelligence. However, prior works typically focused on body behavior, neglecting the crucial role of eye gaze and its synergy with body motion. We…

Human-Computer Interaction · Computer Science 2025-11-21 Qing Chang , Zhiming Hu

When reading, we often have specific information that interests us in a text. For example, you might be reading this paper because you are curious about LLMs for eye movements in reading, the experimental design, or perhaps you wonder…

Computation and Language · Computer Science 2026-03-03 Cfir Avraham Hadar , Omer Shubi , Yoav Meiri , Amit Heshes , Yevgeni Berzak

With the rise in popularity of smart glasses, users' attention has been integrated into Vision-Language Models (VLMs) to streamline multi-modal querying in daily scenarios. However, leveraging gaze data to model users' attention may…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Zeyu Wang , Baiyu Chen , Kun Yan , Hongjing Piao , Hao Xue , Flora D. Salim , Yuanchun Shi , Yuntao Wang

Vision Language Models (VLMs) have demonstrated strong capabilities in understanding visual content, yet their ability to predict where humans look on user interfaces remains unexplored. We present UIGaze, a study investigating how closely…

Human-Computer Interaction · Computer Science 2026-04-30 Min Song , Yoonseong Lee , Yeonhu Seo

Transcripts displayed on dictation interfaces can be hard to read due to recognition errors and disfluencies. LLM-based text auto-correction could help, but changing the text during production could lead to distraction and unintended…

Human-Computer Interaction · Computer Science 2026-03-24 Zhaohui Liang , Yonglin Chen , Naser Al Madi , Can Liu
‹ Prev 1 2 3 10 Next ›