English
Related papers

Related papers: Predicting Gaze in Egocentric Video by Learning Ta…

200 papers

Many computer vision tasks rely on labeled data. Rapid progress in generative modeling has led to the ability to synthesize photorealistic images. However, controlling specific aspects of the generation process such that the data can be…

Computer Vision and Pattern Recognition · Computer Science 2020-10-26 Yufeng Zheng , Seonwook Park , Xucong Zhang , Shalini De Mello , Otmar Hilliges

In the evolving landscape of human-autonomy teaming (HAT), fostering effective collaboration and trust between human and autonomous agents is increasingly important. To explore this, we used the game Overcooked AI to create dynamic teaming…

Human-Computer Interaction · Computer Science 2025-06-18 Anthony J. Ries , Stéphane Aroca-Ouellette , Alessandro Roncone , Ewart J. de Visser

Current LLM assistants are powerful at answering questions, but they have limited access to the behavioral context that reveals when and where a user is struggling. We present a gaze-grounded multimodal LLM assistant that uses egocentric…

Human-Computer Interaction · Computer Science 2026-04-10 Valdemar Danry , Javier Hernandez , Andrew Wilson , Pattie Maes , Judith Amores

Ultra-long egocentric videos spanning multiple days present significant challenges for video understanding. Existing approaches still rely on fragmented local processing and limited temporal modeling, restricting their ability to reason…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Shitong Sun , Ke Han , Yukai Huang , Weitong Cai , Jifei Song

Non-invasive gaze estimation methods usually regress gaze directions directly from a single face or eye image. However, due to important variabilities in eye shapes and inner eye structures amongst individuals, universal models obtain…

Computer Vision and Pattern Recognition · Computer Science 2020-02-07 Gang Liu , Yu Yu , Kenneth A. Funes Mora , Jean-Marc Odobez

To further advance driver monitoring and assistance systems, it is important to understand how drivers allocate their attention, in other words, where do they tend to look and why. Traditionally, factors affecting human visual attention…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Iuliia Kotseruba , John K. Tsotsos

Autonomous driving is a multi-task problem requiring a deep understanding of the visual environment. End-to-end autonomous systems have attracted increasing interest as a method of learning to drive without exhaustively programming…

Computer Vision and Pattern Recognition · Computer Science 2019-09-12 Alexander Makrigiorgos , Ali Shafti , Alex Harston , Julien Gerard , A. Aldo Faisal

Many cameras implement auto-focus functionality. However, they typically require the user to manually identify the location to be focused on. While such an approach works for temporally-sparse autofocusing functionality (e.g., photo…

Computer Vision and Pattern Recognition · Computer Science 2017-11-10 Wolfgang Fuhl , Thiago Santini , Enkelejda Kasneci

Visual saliency patterns are the result of a variety of factors aside from the image being parsed, however existing approaches have ignored these. To address this limitation, we propose a novel saliency estimation model which leverages the…

Computer Vision and Pattern Recognition · Computer Science 2018-03-12 Tharindu Fernando , Simon Denman , Sridha Sridharan , Clinton Fookes

Automatic video captioning is challenging due to the complex interactions in dynamic real scenes. A comprehensive system would ultimately localize and track the objects, actions and interactions present in a video and generate a description…

Computer Vision and Pattern Recognition · Computer Science 2016-10-19 Mihai Zanfir , Elisabeta Marinoiu , Cristian Sminchisescu

For reliable autonomous robot navigation in urban settings, the robot must have the ability to identify semantically traversable terrains in the image based on the semantic understanding of the scene. This reasoning ability is based on…

Robotics · Computer Science 2024-12-30 Yunho Kim , Jeong Hyun Lee , Choongin Lee , Juhyeok Mun , Donghoon Youm , Jeongsoo Park , Jemin Hwangbo

The eye fixation patterns of human observers are a fundamental indicator of the aspects of an image to which humans attend. Thus, manipulating fixation patterns to guide human attention is an exciting challenge in digital image processing.…

Computer Vision and Pattern Recognition · Computer Science 2017-12-19 Leon A. Gatys , Matthias Kümmerer , Thomas S. A. Wallis , Matthias Bethge

Currently successful methods for video description are based on encoder-decoder sentence generation using recur-rent neural networks (RNNs). Recent work has shown the advantage of integrating temporal and/or spatial attention mechanisms…

Computer Vision and Pattern Recognition · Computer Science 2017-03-13 Chiori Hori , Takaaki Hori , Teng-Yok Lee , Kazuhiro Sumi , John R. Hershey , Tim K. Marks

Egocentric videos capture scenes from a wearer's viewpoint, resulting in dynamic backgrounds, frequent motion, and occlusions, posing challenges to accurate keystep recognition. We propose a flexible graph-learning framework for…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Julia Lee Romero , Kyle Min , Subarna Tripathi , Morteza Karimzadeh

Predicting the target of visual search from eye fixation (gaze) data is a challenging problem with many applications in human-computer interaction. In contrast to previous work that has focused on individual instances as a search target, we…

Computer Vision and Pattern Recognition · Computer Science 2017-04-04 Hosnieh Sattar , Andreas Bulling , Mario Fritz

We introduce a multi-stage framework that uses mean curvature on a hand surface and focuses on learning interaction between hand and object by analyzing hand grasp type for hand action recognition in egocentric videos. The proposed method…

Computer Vision and Pattern Recognition · Computer Science 2021-09-09 Sangpil Kim , Jihyun Bae , Hyunggun Chi , Sunghee Hong , Byoung Soo Koh , Karthik Ramani

Understanding and predicting human visuomotor coordination is crucial for applications in robotics, human-computer interaction, and assistive technologies. This work introduces a forecasting-based task for visuomotor modeling, where the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Wenqi Jia , Bolin Lai , Miao Liu , Danfei Xu , James M. Rehg

Graph-based next-step prediction models have recently been very successful in modeling complex high-dimensional physical systems on irregular meshes. However, due to their short temporal attention span, these models suffer from error…

Machine Learning · Computer Science 2022-05-27 Xu Han , Han Gao , Tobias Pfaff , Jian-Xun Wang , Li-Ping Liu

Understanding the human-object interactions (HOIs) from a video is essential to fully comprehend a visual scene. This line of research has been addressed by detecting HOIs from images and lately from videos. However, the video-based HOI…

Computer Vision and Pattern Recognition · Computer Science 2023-06-07 Zhifan Ni , Esteve Valls Mascaró , Hyemin Ahn , Dongheui Lee

Event boundaries play a crucial role as a pre-processing step for detection, localization, and recognition tasks of human activities in videos. Typically, although their intrinsic subjectiveness, temporal bounds are provided manually as…

Computer Vision and Pattern Recognition · Computer Science 2018-09-07 Alejandro Cartas , Estefania Talavera , Petia Radeva , Mariella Dimiccoli