中文
相关论文

相关论文: Spatial Cognition from Egocentric Video: Out of Si…

200 篇论文

Online continual learning from data streams in dynamic environments is a critical direction in the computer vision field. However, realistic benchmarks and fundamental studies in this line are still missing. To bridge the gap, we present a…

计算机视觉与模式识别 · 计算机科学 2021-09-09 Jianren Wang , Xin Wang , Yue Shang-Guan , Abhinav Gupta

Humans possess the visual-spatial intelligence to remember spaces from sequential visual observations. However, can Multimodal Large Language Models (MLLMs) trained on million-scale video datasets also ``think in space'' from videos? We…

计算机视觉与模式识别 · 计算机科学 2025-07-04 Jihan Yang , Shusheng Yang , Anjali W. Gupta , Rilyn Han , Li Fei-Fei , Saining Xie

Perceiving the world from both egocentric (first-person) and exocentric (third-person) perspectives is fundamental to human cognition, enabling rich and complementary understanding of dynamic environments. In recent years, allowing the…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Yuping He , Yifei Huang , Guo Chen , Lidong Lu , Baoqi Pei , Jilan Xu , Tong Lu , Yoichi Sato

Egocentric action anticipation consists in understanding which objects the camera wearer will interact with in the near future and which actions they will perform. We tackle the problem proposing an architecture able to anticipate actions…

计算机视觉与模式识别 · 计算机科学 2019-08-07 Antonino Furnari , Giovanni Maria Farinella

We present EgoFun3D, a coordinated task formulation, dataset, and benchmark for modeling interactive 3D objects from egocentric videos. Interactive objects are of high interest for embodied AI but scarce, making modeling from readily…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Weikun Peng , Denys Iliash , Manolis Savva

Perceiving the physical world in 3D is fundamental for self-driving applications. Although temporal motion is an invaluable resource to human vision for detection, tracking, and depth perception, such features have not been thoroughly…

计算机视觉与模式识别 · 计算机科学 2020-07-21 Garrick Brazil , Gerard Pons-Moll , Xiaoming Liu , Bernt Schiele

For humans, object detection, recognition, and tracking are innate. These provide the ability for human to perceive their environment and objects within their environment. This ability however doesn't translate well in computers. In…

计算机视觉与模式识别 · 计算机科学 2022-03-23 Shiyao Chen , Dale Chen-Song

Human actions in egocentric videos are often hand-object interactions composed from a verb (performed by the hand) applied to an object. Despite their extensive scaling up, egocentric datasets still face two limitations - sparsity of action…

计算机视觉与模式识别 · 计算机科学 2023-12-13 Dibyadip Chatterjee , Fadime Sener , Shugao Ma , Angela Yao

The scale and diversity of demonstration data required for imitation learning is a significant challenge. We present EgoMimic, a full-stack framework which scales manipulation via human embodiment data, specifically egocentric human videos…

机器人学 · 计算机科学 2024-11-01 Simar Kareer , Dhruv Patel , Ryan Punamiya , Pranay Mathur , Shuo Cheng , Chen Wang , Judy Hoffman , Danfei Xu

In this paper we address the problems of detecting objects of interest in a video and of estimating their locations, solely from the gaze directions of people present in the video. Objects can be indistinctly located inside or outside the…

计算机视觉与模式识别 · 计算机科学 2019-03-01 Benoit Massé , Stéphane Lathuilière , Pablo Mesejo , Radu Horaud

Recent advances in Multi-modal Large Language Models (MLLMs) have showcased remarkable capabilities in vision-language understanding. However, enabling robust video spatial reasoning-the ability to comprehend object locations, orientations,…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Haoran Tang , Meng Cao , Ruyang Liu , Xiaoxi Liang , Linglong Li , Ge Li , Xiaodan Liang

Unlike traditional third-person cameras mounted on robots, a first-person camera, captures a person's visual sensorimotor object interactions from up close. In this paper, we study the tight interplay between our momentary visual attention…

计算机视觉与模式识别 · 计算机科学 2017-06-13 Gedas Bertasius , Hyun Soo Park , Stella X. Yu , Jianbo Shi

World-wide detailed 2D maps require enormous collective efforts. OpenStreetMap is the result of 11 million registered users manually annotating the GPS location of over 1.75 billion entries, including distinctive landmarks and common urban…

计算机视觉与模式识别 · 计算机科学 2024-11-25 Matteo Toso , Stefano Fiorini , Stuart James , Alessio Del Bue

Egocentric sensors such as AR/VR devices capture human-object interactions and offer the potential to provide task-assistance by recalling 3D locations of objects of interest in the surrounding environment. This capability requires instance…

计算机视觉与模式识别 · 计算机科学 2024-06-10 Yunhan Zhao , Haoyu Ma , Shu Kong , Charless Fowlkes

From Vision-Language-Action (VLA) systems to robotics, existing egocentric datasets primarily focus on action recognition tasks, while largely overlooking the inherent role of motion analysis in sports and other fast-movement scenarios. To…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Si-En Hong , James Tribble , Alexander Lake , Hao Wang , Chaoyi Zhou , Ashish Bastola , Siyu Huang , Eisa Chaudhary , Brian Canada , Ismahan Arslan-Ari , Abolfazl Razi

An important challenge for autonomous agents such as robots is to maintain a spatially and temporally consistent model of the world. It must be maintained through occlusions, previously-unseen views, and long time horizons (e.g., loop…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Dominik A. Kloepfer , Dylan Campbell , João F. Henriques

Visual queries 3D localization (VQ3D) is a task in the Ego4D Episodic Memory Benchmark. Given an egocentric video, the goal is to answer queries of the form "Where did I last see object X?", where the query object X is specified as a static…

计算机视觉与模式识别 · 计算机科学 2022-11-21 Jinjie Mai , Chen Zhao , Abdullah Hamdi , Silvio Giancola , Bernard Ghanem

We interact with the world with our hands and see it through our own (egocentric) perspective. A holistic 3Dunderstanding of such interactions from egocentric views is important for tasks in robotics, AR/VR, action recognition and motion…

Spatiotemporal video grounding aims to localize target entities in videos based on textual queries. While existing research has made significant progress in exocentric videos, the egocentric setting remains relatively underexplored, despite…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Shuo Liang , Yiwu Zhong , Zi-Yuan Hu , Yeyao Tao , Liwei Wang

3D object tracking is a critical task in autonomous driving systems. It plays an essential role for the system's awareness about the surrounding environment. At the same time there is an increasing interest in algorithms for autonomous cars…

计算机视觉与模式识别 · 计算机科学 2022-10-31 Nicola Marinello , Marc Proesmans , Luc Van Gool