中文
相关论文

相关论文: EGO-TOPO: Environment Affordances from Egocentric …

200 篇论文

Recent technological advances have made lightweight, head mounted cameras both practical and affordable and products like Google Glass show first approaches to introduce the idea of egocentric (first-person) video to the mainstream.…

计算机视觉与模式识别 · 计算机科学 2015-01-14 Sven Bambach

In a wearable camera video, we see what the camera wearer sees. While this makes it easy to know roughly what he chose to look at, it does not immediately reveal when he was engaged with the environment. Specifically, at what moments did…

计算机视觉与模式识别 · 计算机科学 2016-04-05 Yu-Chuan Su , Kristen Grauman

Human actions in egocentric videos are often hand-object interactions composed from a verb (performed by the hand) applied to an object. Despite their extensive scaling up, egocentric datasets still face two limitations - sparsity of action…

计算机视觉与模式识别 · 计算机科学 2023-12-13 Dibyadip Chatterjee , Fadime Sener , Shugao Ma , Angela Yao

Robots need to understand their environment to perform their task. If it is possible to pre-program a visual scene analysis process in closed environments, robots operating in an open environment would benefit from the ability to learn it…

机器人学 · 计算机科学 2019-03-12 Leni K. Le Goff , Oussama Yaakoubi , Alexandre Coninx , Stephane Doncieux

Advancements in egocentric video datasets like Ego4D, EPIC-Kitchens, and Ego-Exo4D have enriched the study of first-person human interactions, which is crucial for applications in augmented reality and assisted living. Despite these…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Joungbin An , Yunsu Park , Hyolim Kang , Seon Joo Kim

Spatiotemporal video grounding aims to localize target entities in videos based on textual queries. While existing research has made significant progress in exocentric videos, the egocentric setting remains relatively underexplored, despite…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Shuo Liang , Yiwu Zhong , Zi-Yuan Hu , Yeyao Tao , Liwei Wang

With the surge in attention to Egocentric Hand-Object Interaction (Ego-HOI), large-scale datasets such as Ego4D and EPIC-KITCHENS have been proposed. However, most current research is built on resources derived from third-person video…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Yue Xu , Yong-Lu Li , Zhemin Huang , Michael Xu Liu , Cewu Lu , Yu-Wing Tai , Chi-Keung Tang

Although First Person Vision systems can sense the environment from the user's perspective, they are generally unable to predict his intentions and goals. Since human activities can be decomposed in terms of atomic actions and interactions…

计算机视觉与模式识别 · 计算机科学 2019-04-11 Antonino Furnari , Sebastiano Battiato , Kristen Grauman , Giovanni Maria Farinella

Monocular egocentric human pose estimation is essential for ubiquitous activity monitoring. However, understanding the user's absolute location within the environment remains a challenge. Existing methods primarily focus on relative motion…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Hiroyuki Deguchi , Ryosuke Hori , Kotaro Amaya , Tsubasa Maruyama , Mitsunori Tada , Hideo Saito

Recently, there has been a growing interest in wearable sensors which provides new research perspectives for 360 {\deg} video analysis. However, the lack of 360 {\deg} datasets in literature hinders the research in this field. To bridge…

计算机视觉与模式识别 · 计算机科学 2020-10-19 Keshav Bhandari , Mario A. DeLaGarza , Ziliang Zong , Hugo Latapie , Yan Yan

Affordance grounding aims to locate objects' "action possibilities" regions, which is an essential step toward embodied intelligence. Due to the diversity of interactive affordance, the uniqueness of different individuals leads to diverse…

计算机视觉与模式识别 · 计算机科学 2023-05-26 Hongchen Luo , Wei Zhai , Jing Zhang , Yang Cao , Dacheng Tao

Human activities exhibit a strong correlation between actions and the places where these are performed, such as washing something at a sink. More specifically, in daily living environments we may identify particular locations, hereinafter…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Simone Alberto Peirone , Gabriele Goletto , Mirco Planamente , Andrea Bottino , Barbara Caputo , Giuseppe Averta

When people observe and interact with physical spaces, they are able to associate functionality to regions in the environment. Our goal is to automate dense functional understanding of large spaces by leveraging sparse activity…

计算机视觉与模式识别 · 计算机科学 2016-05-06 Nicholas Rhinehart , Kris M. Kitani

How do we know that a kitchen is a kitchen by looking? Relatively little is known about how we conceptualize and categorize different visual environments. Traditional models of visual perception posit that scene categorization is achieved…

神经元与认知 · 定量生物学 2014-11-20 Michelle R. Greene , Christopher Baldassano , Andre Esteva , Diane M. Beck , Li Fei-Fei

We investigate exocentric-to-egocentric cross-view translation, which aims to generate a first-person (egocentric) view of an actor based on a video recording that captures the actor from a third-person (exocentric) perspective. To this…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Mi Luo , Zihui Xue , Alex Dimakis , Kristen Grauman

Short Term object-interaction Anticipation consists in detecting the location of the next active objects, the noun and verb categories of the interaction, as well as the time to contact from the observation of egocentric video. This ability…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Lorenzo Mur Labadia , Ruben Martinez-Cantin , Jose J. Guerrero , Giovanni M. Farinella , Antonino Furnari

Short-term action anticipation (STA) in first-person videos is a challenging task that involves understanding the next active object interactions and predicting future actions. Existing action anticipation methods have primarily focused on…

计算机视觉与模式识别 · 计算机科学 2023-06-26 Sanket Thakur , Cigdem Beyan , Pietro Morerio , Vittorio Murino , Alessio Del Bue

The body pose of a person wearing a camera is of great interest for applications in augmented reality, healthcare, and robotics, yet much of the person's body is out of view for a typical wearable camera. We propose a learning-based…

计算机视觉与模式识别 · 计算机科学 2020-03-31 Evonne Ng , Donglai Xiang , Hanbyul Joo , Kristen Grauman

In this paper, we present an approach for robot learning of social affordance from human activity videos. We consider the problem in the context of human-robot interaction: Our approach learns structural representations of human-human (and…

机器人学 · 计算机科学 2016-04-22 Tianmin Shu , M. S. Ryoo , Song-Chun Zhu

Egocentric videos are characterised by their ability to have the first person view. With the popularity of Google Glass and GoPro, use of egocentric videos is on the rise. Recognizing action of the wearer from egocentric videos is an…

计算机视觉与模式识别 · 计算机科学 2016-04-08 Suriya Singh , Chetan Arora , C. V. Jawahar