中文
相关论文

相关论文: Gaze-Regularized VLMs for Ego-Centric Behavior Und…

200 篇论文

In this paper, we present the first transformer-based model to address the challenging problem of egocentric gaze estimation. We observe that the connection between the global scene context and local visual information is vital for…

计算机视觉与模式识别 · 计算机科学 2024-10-17 Bolin Lai , Miao Liu , Fiona Ryan , James M. Rehg

People continuously perceive and interact with their surroundings based on underlying intentions that drive their exploration and behaviors. While research in egocentric user and scene understanding has focused primarily on motion and…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Heeseung Yun , Joonil Na , Jaeyeon Kim , Calvin Murdock , Gunhee Kim

Understanding human social behavior such as recognizing emotions and the social dynamics causing them is an important and challenging problem. While LLMs have made remarkable advances, they are limited to the textual domain and cannot…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Tania Chakraborty , Eylon Caplan , Dan Goldwasser

With the rise in popularity of smart glasses, users' attention has been integrated into Vision-Language Models (VLMs) to streamline multi-modal querying in daily scenarios. However, leveraging gaze data to model users' attention may…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Zeyu Wang , Baiyu Chen , Kun Yan , Hongjing Piao , Hao Xue , Flora D. Salim , Yuanchun Shi , Yuntao Wang

Eye gaze analysis is an important research problem in the field of Computer Vision and Human-Computer Interaction. Even with notable progress in the last 10 years, automatic gaze analysis still remains challenging due to the uniqueness of…

计算机视觉与模式识别 · 计算机科学 2022-07-22 Shreya Ghosh , Abhinav Dhall , Munawar Hayat , Jarrod Knibbe , Qiang Ji

A considerable part of the performance of today's large language models (LLM's) and multimodal large language models (MLLM's) depends on their tokenization strategies. While tokenizers are extensively researched for textual and visual…

机器学习 · 计算机科学 2025-03-31 Tim Rolff , Jurik Karimian , Niklas Hypki , Susanne Schmidt , Markus Lappe , Frank Steinicke

Personalized driving refers to an autonomous vehicle's ability to adapt its driving behavior or control strategies to match individual users' preferences and driving styles while maintaining safety and comfort standards. However, existing…

Driver visual attention prediction is a critical task in autonomous driving and human-computer interaction (HCI) research. Most prior studies focus on estimating attention allocation at a single moment in time, typically using static RGB…

计算机视觉与模式识别 · 计算机科学 2025-08-11 Kaiser Hamid , Khandakar Ashrafi Akbar , Nade Liang

Understanding human activity is a crucial yet intricate task in egocentric vision, a field that focuses on capturing visual perspectives from the camera wearer's viewpoint. Traditional methods heavily rely on representation learning that is…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Sanghwan Kim , Daoji Huang , Yongqin Xian , Otmar Hilliges , Luc Van Gool , Xi Wang

Online egocentric gaze estimation predicts where a camera wearer is looking from first-person video using only past and current frames, a task essential for augmented reality and assistive technologies. Unlike third-person gaze estimation,…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Jia Li , Wenjie Zhao , Shijian Deng , Bolin Lai , Yuheng Wu , RUijia Chen , Jon E. Froehlich , Yuhang Zhao , Yapeng Tian

Designing intuitive interfaces for robotic control remains a central challenge in enabling effective human-robot interaction, particularly in assistive care settings. Eye gaze offers a fast, non-intrusive, and intent-rich input modality,…

机器人学 · 计算机科学 2026-01-12 Tracey Yee Hsin Tay , Xu Yan , Jonathan Ouyang , Daniel Wu , William Jiang , Jonathan Kao , Yuchen Cui

Gaze following and social gaze prediction are fundamental tasks providing insights into human communication behaviors, intent, and social interactions. Most previous approaches addressed these tasks separately, either by designing highly…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Anshul Gupta , Samy Tafasca , Arya Farkhondeh , Pierre Vuillecard , Jean-Marc Odobez

Technical progress in hardware and software enables us to record gaze data in everyday situations and over long time spans. Among a multitude of research opportunities, this technology enables visualization researchers to catch a glimpse…

人机交互 · 计算机科学 2023-12-05 Maurice Koch , Kuno Kurzhals , Michael Burch , Daniel Weiskopf

Vision-language models (VLMs) have made substantial progress across a wide range of visual question answering benchmarks, spanning visual reasoning, document understanding, and multimodal dialogue. These improvements are evident in a wide…

计算机视觉与模式识别 · 计算机科学 2026-02-23 Dhruba Ghosh , Yuhui Zhang , Ludwig Schmidt

Understanding human gaze behavior is essential for complex scene comprehension and human-computer interaction. Traditional gaze following models are typically restricted to pure spatial localization, lacking the high-level capacity to…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Qiaomu Miao , Haoyu Wu , Jingyi Xu , Minh Hoai , Dimitris Samaras

Smart glasses with AI assistants are increasingly used in daily life. However, current systems lack awareness of the user's internal cognitive state, leaving them unable to proactively anticipate users' needs without access to cognitive…

Code summarization is the task of generating natural language descriptions of source code, which is critical for software comprehension and maintenance. While large language models (LLMs) have achieved remarkable progress on this task, an…

软件工程 · 计算机科学 2026-02-27 Jiahao Zhang , Yifan Zhang , Kevin Leach , Yu Huang

Humans perceive their visual environment by directing their eyes towards relevant objects. The deployment of visual attention depends substantially on the stimulus's properties, higher cognitive processes, and biases and constraints of the…

神经元与认知 · 定量生物学 2025-07-22 Thomas Fabian

Eye-tracking data reveals valuable insights into users' cognitive states but is difficult to analyze due to its structured, non-linguistic nature. While large language models (LLMs) excel at reasoning over text, they struggle with temporal…

人机交互 · 计算机科学 2025-07-25 Dongyang Guo , Yasmeen Abdrabou , Enkeleda Thaqi , Enkelejda Kasneci

This paper addresses the challenging problem of estimating the general visual attention of people in images. Our proposed method is designed to work across multiple naturalistic social scenarios and provides a full picture of the subject's…

计算机视觉与模式识别 · 计算机科学 2018-07-30 Eunji Chong , Nataniel Ruiz , Yongxin Wang , Yun Zhang , Agata Rozga , James Rehg