中文
相关论文

相关论文: Robust Egocentric Visual Attention Prediction Thro…

200 篇论文

Active perception, the ability of a robot to proactively adjust its viewpoint to acquire task-relevant information, is essential for robust operation in unstructured real-world environments. While critical for downstream tasks such as…

机器人学 · 计算机科学 2026-03-03 Yongxi Huang , Zhuohang Wang , Wenjing Tang , Cewu Lu , Panpan Cai

We train models to Predict Ego-centric Video from human Actions (PEVA), given the past video and an action represented by the relative 3D body pose. By conditioning on kinematic pose trajectories, structured by the joint hierarchy of the…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Yutong Bai , Danny Tran , Amir Bar , Yann LeCun , Trevor Darrell , Jitendra Malik

Modeling the cognitive and experiential factors of human navigation is central to deepening our understanding of human-environment interaction and to enabling safe social navigation and effective assistive wayfinding. Most existing methods…

机器学习 · 计算机科学 2026-03-09 Zhiwen Qiu , Ziang Liu , Wenqian Niu , Tapomayukh Bhattacharjee , Saleh Kalantari

Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work, we propose EgoControl, a pose-controllable video diffusion…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Enrico Pallotta , Sina Mokhtarzadeh Azar , Lars Doorenbos , Serdar Ozsoy , Umar Iqbal , Juergen Gall

Given a video captured from a first person perspective and the environment context of where the video is recorded, can we recognize what the person is doing and identify where the action occurs in the 3D space? We address this challenging…

计算机视觉与模式识别 · 计算机科学 2022-08-16 Miao Liu , Lingni Ma , Kiran Somasundaram , Yin Li , Kristen Grauman , James M. Rehg , Chao Li

Inspired by recent work in machine translation and object detection, we introduce an attention based model that automatically learns to describe the content of images. We describe how we can train this model in a deterministic manner using…

Both humans and machines learn the meaning of unknown words through contextual information in a sentence, but not all contexts are equally helpful for learning. We introduce an effective method for capturing the level of contextual…

计算与语言 · 计算机科学 2023-11-10 Sungjin Nam , David Jurgens , Gwen Frishkoff , Kevyn Collins-Thompson

Egocentric gaze anticipation serves as a key building block for the emerging capability of Augmented Reality. Notably, gaze behavior is driven by both visual cues and audio signals during daily activities. Motivated by this observation, we…

计算机视觉与模式识别 · 计算机科学 2024-03-25 Bolin Lai , Fiona Ryan , Wenqi Jia , Miao Liu , James M. Rehg

Nowadays, there is an upsurge of interest in using lifelogging devices. Such devices generate huge amounts of image data; consequently, the need for automatic methods for analyzing and summarizing these data is drastically increasing. We…

计算机视觉与模式识别 · 计算机科学 2019-05-13 Estefania Talavera , Nicolai Petkov , Petia Radeva

Object-centric representations are a promising path toward more systematic generalization by providing flexible abstractions upon which compositional world models can be built. Recent work on simple 2D and 3D datasets has shown that models…

Robust and efficient learning remains a challenging problem in robotics, in particular with complex visual inputs. Inspired by human attention mechanism, with which we quickly process complex visual scenes and react to changes in the…

机器人学 · 计算机科学 2023-08-30 Daniel Scheuchenstuhl , Stefan Ulmer , Felix Resch , Luigi Berducci , Radu Grosu

In this work, we tackle two vital tasks in automated driving systems, i.e., driver intent prediction and risk object identification from egocentric images. Mainly, we investigate the question: what would be good road scene-level…

计算机视觉与模式识别 · 计算机科学 2023-03-01 Zihao Xiao , Alan Yuille , Yi-Ting Chen

Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for understanding these…

计算机视觉与模式识别 · 计算机科学 2022-01-07 Hao Jiang , Calvin Murdock , Vamsi Krishna Ithapu

This paper presents an unsupervised approach towards automatically extracting video-based guidance on object usage, from egocentric video and wearable gaze tracking, collected from multiple users while performing tasks. The approach i)…

计算机视觉与模式识别 · 计算机科学 2016-03-22 Dima Damen , Teesid Leelasawassuk , Walterio Mayol-Cuevas

Communicating in noisy, multi-talker environments is challenging, especially for people with hearing impairments. Egocentric video data can potentially be used to identify a user's conversation partners, which could be used to inform…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Tobias Dorszewski , Søren A. Fuglsang , Jens Hjortkjær

Prediction of dynamic environment is crucial to safe navigation of an autonomous vehicle. Urban traffic scenes are particularly challenging to forecast due to complex interactions between various dynamic agents, such as vehicles and…

计算机视觉与模式识别 · 计算机科学 2023-01-12 Rabbia Asghar , Lukas Rummelhard , Anne Spalanzani , Christian Laugier

In this paper, we present the first transformer-based model to address the challenging problem of egocentric gaze estimation. We observe that the connection between the global scene context and local visual information is vital for…

计算机视觉与模式识别 · 计算机科学 2024-10-17 Bolin Lai , Miao Liu , Fiona Ryan , James M. Rehg

Anticipating human actions is an important task that needs to be addressed for the development of reliable intelligent agents, such as self-driving cars or robot assistants. While the ability to make future predictions with high accuracy is…

计算机视觉与模式识别 · 计算机科学 2021-07-21 Olga Zatsarynna , Yazan Abu Farha , Juergen Gall

We seek to accelerate research in developing rich, multimodal scene models trained from egocentric data, based on differentiable volumetric ray-tracing inspired by Neural Radiance Fields (NeRFs). The construction of a NeRF-like model from…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Jiankai Sun , Jianing Qiu , Chuanyang Zheng , John Tucker , Javier Yu , Mac Schwager

Egocentric perception enables humans to experience and understand the world directly from their own point of view. Translating exocentric (third-person) videos into egocentric (first-person) videos opens up new possibilities for immersive…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Taewoong Kang , Kinam Kim , Dohyeon Kim , Minho Park , Junha Hyung , Jaegul Choo