English
Related papers

Related papers: EgoFun3D: Modeling Interactive Objects from Egocen…

200 papers

This paper deals with the problem of localizing objects in image and video datasets from visual exemplars. In particular, we focus on the challenging problem of egocentric visual query localization. We first identify grave implicit biases…

Computer Vision and Pattern Recognition · Computer Science 2023-04-07 Mengmeng Xu , Yanghao Li , Cheng-Yang Fu , Bernard Ghanem , Tao Xiang , Juan-Manuel Perez-Rua

Large-scale egocentric video datasets capture diverse human activities across a wide range of scenarios, offering rich and detailed insights into how humans interact with objects, especially those that require fine-grained dexterous…

The concept of 3D scene graphs is increasingly recognized as a powerful semantic and hierarchical representation of the environment. Current approaches often address this at a coarse, object-level resolution. In contrast, our goal is to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Dennis Rotondi , Fabio Scaparro , Hermann Blum , Kai O. Arras

What if a video generation model could not only imagine a plausible future, but the correct one, accurately reflecting how the world changes with each action? We address this question by presenting the Egocentric World Model (EgoWM), a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Anurag Bagchi , Zhipeng Bao , Homanga Bharadhwaj , Yu-Xiong Wang , Pavel Tokmakov , Martial Hebert

Computer vision is largely based on 2D techniques, with 3D vision still relegated to a relatively narrow subset of applications. However, by building on recent advances in 3D models such as neural radiance fields, some authors have shown…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Vadim Tschernezki , Diane Larlus , Iro Laina , Andrea Vedaldi

Being able to map the activities of others into one's own point of view is one fundamental human skill even from a very early age. Taking a step toward understanding this human ability, we introduce EgoExoLearn, a large-scale dataset that…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Yifei Huang , Guo Chen , Jilan Xu , Mingfang Zhang , Lijin Yang , Baoqi Pei , Hongjie Zhang , Lu Dong , Yali Wang , Limin Wang , Yu Qiao

Wearable cameras allow to collect images and videos of humans interacting with the world. While human-object interactions have been thoroughly investigated in third person vision, the problem has been understudied in egocentric settings and…

Computer Vision and Pattern Recognition · Computer Science 2020-10-13 Francesco Ragusa , Antonino Furnari , Salvatore Livatino , Giovanni Maria Farinella

Egocentric videos can bring a lot of information about how humans perceive the world and interact with the environment, which can be beneficial for the analysis of human behaviour. The research in egocentric video analysis is developing…

Computer Vision and Pattern Recognition · Computer Science 2021-07-29 Ivan Rodin , Antonino Furnari , Dimitrios Mavroedis , Giovanni Maria Farinella

Wearable cameras allow to acquire images and videos from the user's perspective. These data can be processed to understand humans behavior. Despite human behavior analysis has been thoroughly investigated in third person vision, it is still…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Francesco Ragusa , Antonino Furnari , Giovanni Maria Farinella

In this technical report, we present our solution for the EgoPlan Challenge in ICML 2024. To address the real-world egocentric task planning problem, we introduce a novel planning framework which comprises three stages: long-term memory…

Robotics · Computer Science 2024-07-30 Letian Shi , Qi Lv , Xiang Deng , Liqiang Nie

In this paper, we propose a novel approach to enhance the 3D body pose estimation of a person computed from videos captured from a single wearable camera. The key idea is to leverage high-level features linking first- and third-views in a…

Computer Vision and Pattern Recognition · Computer Science 2022-06-16 Ameya Dhamanaskar , Mariella Dimiccoli , Enric Corona , Albert Pumarola , Francesc Moreno-Noguer

Recent Multimodal Large Language Models (MLLMs) have shown high potential for spatial reasoning within 3D scenes. However, they typically rely on computationally expensive 3D representations like point clouds or reconstructed Bird's-Eye…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Shuyao Shi , Kang G. Shin

In this paper, we address the challenge of understanding human activities from an egocentric perspective. Traditional activity recognition techniques face unique challenges in egocentric videos due to the highly dynamic nature of the head…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Zachary Chavis , Stephen J. Guy , Hyun Soo Park

Accurate 3D understanding of human hands and objects during manipulation remains a significant challenge for egocentric computer vision. Existing hand-object interaction datasets are predominantly captured in controlled studio settings,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Patrick Rim , Kevin Harris , Braden Copple , Shangchen Han , Xu Xie , Ivan Shugurov , Sizhe An , He Wen , Alex Wong , Tomas Hodan , Kun He

Articulated objects are fundamental for robotics, simulation of physics, and interactive virtual environments. However, reconstructing them from visual input remains challenging, as it requires jointly inferring both part geometry and…

Robotics · Computer Science 2026-03-17 Zhuangzhe Wu , Yue Xin , Chengkai Hou , Minghao Chen , Yaoxu Lyu , Jieyu Zhang , Shanghang Zhang

Tactile recognition of 3D objects remains a challenging task. Compared to 2D shapes, the complex geometry of 3D surfaces requires richer tactile signals, more dexterous actions, and more advanced encoding techniques. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2023-03-01 Jingxi Xu , Han Lin , Shuran Song , Matei Ciocarlie

First-person video naturally brings the use of a physical environment to the forefront, since it shows the camera wearer interacting fluidly in a space based on his intentions. However, current methods largely separate the observed actions…

Computer Vision and Pattern Recognition · Computer Science 2020-03-31 Tushar Nagarajan , Yanghao Li , Christoph Feichtenhofer , Kristen Grauman

We propose a method for incorporating object interaction and human body dynamics into the task of 3D ego-pose estimation using a head-mounted camera. We use a kinematics model of the human body to represent the entire range of human motion,…

Computer Vision and Pattern Recognition · Computer Science 2020-12-10 Zhengyi Luo , Ryo Hachiuma , Ye Yuan , Shun Iwase , Kris M. Kitani

Egocentric 3D human pose estimation (HPE) from images is challenging due to severe self-occlusions and strong distortion introduced by the fish-eye view from the head mounted camera. Although existing works use intermediate heatmap-based…

Computer Vision and Pattern Recognition · Computer Science 2022-06-13 Jinman Park , Kimathi Kaai , Saad Hossain , Norikatsu Sumi , Sirisha Rambhatla , Paul Fieguth

Egocentric vision captures the scene from the point of view of the camera wearer, while exocentric vision captures the overall scene context. Jointly modeling ego and exo views is crucial to developing next-generation AI agents. The…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Anirudh Thatipelli , Shao-Yuan Lo , Amit K. Roy-Chowdhury