中文
相关论文

相关论文: Predicting Gaze in Egocentric Video by Learning Ta…

200 篇论文

Many computer vision tasks rely on labeled data. Rapid progress in generative modeling has led to the ability to synthesize photorealistic images. However, controlling specific aspects of the generation process such that the data can be…

计算机视觉与模式识别 · 计算机科学 2020-10-26 Yufeng Zheng , Seonwook Park , Xucong Zhang , Shalini De Mello , Otmar Hilliges

In the evolving landscape of human-autonomy teaming (HAT), fostering effective collaboration and trust between human and autonomous agents is increasingly important. To explore this, we used the game Overcooked AI to create dynamic teaming…

人机交互 · 计算机科学 2025-06-18 Anthony J. Ries , Stéphane Aroca-Ouellette , Alessandro Roncone , Ewart J. de Visser

Current LLM assistants are powerful at answering questions, but they have limited access to the behavioral context that reveals when and where a user is struggling. We present a gaze-grounded multimodal LLM assistant that uses egocentric…

人机交互 · 计算机科学 2026-04-10 Valdemar Danry , Javier Hernandez , Andrew Wilson , Pattie Maes , Judith Amores

Ultra-long egocentric videos spanning multiple days present significant challenges for video understanding. Existing approaches still rely on fragmented local processing and limited temporal modeling, restricting their ability to reason…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Shitong Sun , Ke Han , Yukai Huang , Weitong Cai , Jifei Song

Non-invasive gaze estimation methods usually regress gaze directions directly from a single face or eye image. However, due to important variabilities in eye shapes and inner eye structures amongst individuals, universal models obtain…

计算机视觉与模式识别 · 计算机科学 2020-02-07 Gang Liu , Yu Yu , Kenneth A. Funes Mora , Jean-Marc Odobez

To further advance driver monitoring and assistance systems, it is important to understand how drivers allocate their attention, in other words, where do they tend to look and why. Traditionally, factors affecting human visual attention…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Iuliia Kotseruba , John K. Tsotsos

Autonomous driving is a multi-task problem requiring a deep understanding of the visual environment. End-to-end autonomous systems have attracted increasing interest as a method of learning to drive without exhaustively programming…

计算机视觉与模式识别 · 计算机科学 2019-09-12 Alexander Makrigiorgos , Ali Shafti , Alex Harston , Julien Gerard , A. Aldo Faisal

Many cameras implement auto-focus functionality. However, they typically require the user to manually identify the location to be focused on. While such an approach works for temporally-sparse autofocusing functionality (e.g., photo…

计算机视觉与模式识别 · 计算机科学 2017-11-10 Wolfgang Fuhl , Thiago Santini , Enkelejda Kasneci

Visual saliency patterns are the result of a variety of factors aside from the image being parsed, however existing approaches have ignored these. To address this limitation, we propose a novel saliency estimation model which leverages the…

计算机视觉与模式识别 · 计算机科学 2018-03-12 Tharindu Fernando , Simon Denman , Sridha Sridharan , Clinton Fookes

Automatic video captioning is challenging due to the complex interactions in dynamic real scenes. A comprehensive system would ultimately localize and track the objects, actions and interactions present in a video and generate a description…

计算机视觉与模式识别 · 计算机科学 2016-10-19 Mihai Zanfir , Elisabeta Marinoiu , Cristian Sminchisescu

For reliable autonomous robot navigation in urban settings, the robot must have the ability to identify semantically traversable terrains in the image based on the semantic understanding of the scene. This reasoning ability is based on…

机器人学 · 计算机科学 2024-12-30 Yunho Kim , Jeong Hyun Lee , Choongin Lee , Juhyeok Mun , Donghoon Youm , Jeongsoo Park , Jemin Hwangbo

The eye fixation patterns of human observers are a fundamental indicator of the aspects of an image to which humans attend. Thus, manipulating fixation patterns to guide human attention is an exciting challenge in digital image processing.…

计算机视觉与模式识别 · 计算机科学 2017-12-19 Leon A. Gatys , Matthias Kümmerer , Thomas S. A. Wallis , Matthias Bethge

Currently successful methods for video description are based on encoder-decoder sentence generation using recur-rent neural networks (RNNs). Recent work has shown the advantage of integrating temporal and/or spatial attention mechanisms…

计算机视觉与模式识别 · 计算机科学 2017-03-13 Chiori Hori , Takaaki Hori , Teng-Yok Lee , Kazuhiro Sumi , John R. Hershey , Tim K. Marks

Egocentric videos capture scenes from a wearer's viewpoint, resulting in dynamic backgrounds, frequent motion, and occlusions, posing challenges to accurate keystep recognition. We propose a flexible graph-learning framework for…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Julia Lee Romero , Kyle Min , Subarna Tripathi , Morteza Karimzadeh

Predicting the target of visual search from eye fixation (gaze) data is a challenging problem with many applications in human-computer interaction. In contrast to previous work that has focused on individual instances as a search target, we…

计算机视觉与模式识别 · 计算机科学 2017-04-04 Hosnieh Sattar , Andreas Bulling , Mario Fritz

We introduce a multi-stage framework that uses mean curvature on a hand surface and focuses on learning interaction between hand and object by analyzing hand grasp type for hand action recognition in egocentric videos. The proposed method…

计算机视觉与模式识别 · 计算机科学 2021-09-09 Sangpil Kim , Jihyun Bae , Hyunggun Chi , Sunghee Hong , Byoung Soo Koh , Karthik Ramani

Understanding and predicting human visuomotor coordination is crucial for applications in robotics, human-computer interaction, and assistive technologies. This work introduces a forecasting-based task for visuomotor modeling, where the…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Wenqi Jia , Bolin Lai , Miao Liu , Danfei Xu , James M. Rehg

Graph-based next-step prediction models have recently been very successful in modeling complex high-dimensional physical systems on irregular meshes. However, due to their short temporal attention span, these models suffer from error…

机器学习 · 计算机科学 2022-05-27 Xu Han , Han Gao , Tobias Pfaff , Jian-Xun Wang , Li-Ping Liu

Understanding the human-object interactions (HOIs) from a video is essential to fully comprehend a visual scene. This line of research has been addressed by detecting HOIs from images and lately from videos. However, the video-based HOI…

计算机视觉与模式识别 · 计算机科学 2023-06-07 Zhifan Ni , Esteve Valls Mascaró , Hyemin Ahn , Dongheui Lee

Event boundaries play a crucial role as a pre-processing step for detection, localization, and recognition tasks of human activities in videos. Typically, although their intrinsic subjectiveness, temporal bounds are provided manually as…

计算机视觉与模式识别 · 计算机科学 2018-09-07 Alejandro Cartas , Estefania Talavera , Petia Radeva , Mariella Dimiccoli