English
Related papers

Related papers: Localizing Active Objects from Egocentric Vision w…

200 papers

Active Position Estimation (APE) is the task of localizing one or more targets using one or more sensing platforms. APE is a key task for search and rescue missions, wildlife monitoring, source term estimation, and collaborative mobile…

Robotics · Computer Science 2024-10-28 Luca Varotto , Angelo Cenedese , Andrea Cavallaro

It is challenging for humans -- particularly those living with physical disabilities -- to control high-dimensional, dexterous robots. Prior work explores learning embedding functions that map a human's low-dimensional inputs (e.g., via a…

Robotics · Computer Science 2021-05-04 Siddharth Karamcheti , Albert J. Zhai , Dylan P. Losey , Dorsa Sadigh

This study addresses a task designed to predict the future success or failure of open-vocabulary object manipulation. In this task, the model is required to make predictions based on natural language instructions, egocentric view images…

Robotics · Computer Science 2025-01-09 Motonari Kambara , Komei Sugiura

We study the problem of learning a range of vision-based manipulation tasks from a large offline dataset of robot interaction. In order to accomplish this, humans need easy and effective ways of specifying tasks to the robot. Goal images…

Robotics · Computer Science 2021-11-02 Suraj Nair , Eric Mitchell , Kevin Chen , Brian Ichter , Silvio Savarese , Chelsea Finn

Egocentric video understanding is inherently complex due to the dynamic 4D nature of the environment, where camera motion and object displacements necessitate a continuous re-evaluation of spatial relations. In this work, we target a suite…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Fangrui Zhu , Yunfeng Xi , Jianmo Ni , Mu Cai , Boqing Gong , Long Zhao , Chen Qu , Ian Miao , Yi Li , Cheng Zhong , Huaizu Jiang , Shwetak Patel

To enable progress towards egocentric agents capable of understanding everyday tasks specified in natural language, we propose a benchmark and a synthetic dataset called Egocentric Task Verification (EgoTV). The goal in EgoTV is to verify…

Computer Vision and Pattern Recognition · Computer Science 2025-11-13 Rishi Hazra , Brian Chen , Akshara Rai , Nitin Kamra , Ruta Desai

Egocentric video-language pretraining has significantly advanced video representation learning. Humans perceive and interact with a fully 3D world, developing spatial awareness that extends beyond text-based understanding. However, most…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Boshen Xu , Yuting Mei , Xinbi Liu , Sipeng Zheng , Qin Jin

We focus on the task of language-conditioned object placement, in which a robot should generate placements that satisfy all the spatial relational constraints in language instructions. Previous works based on rule-based language parsing or…

Robotics · Computer Science 2023-04-07 Zhixuan Xu , Kechun Xu , Yue Wang , Rong Xiong

In this work, we introduce (a) the new problem of anticipating object state changes in images and videos during procedural activities, (b) new curated annotation data for object state change classification based on the Ego4D dataset, and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Victoria Manousaki , Konstantinos Bacharidis , Filippos Gouidis , Konstantinos Papoutsakis , Dimitris Plexousakis , Antonis Argyros

For effective interactions with the open world, robots should understand how interactions with known and novel objects help them towards their goal. A key aspect of this understanding lies in detecting an object's affordances, which…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Anne Kemmeren , Gertjan Burghouts , Michael van Bekkum , Wouter Meijer , Jelle van Mil

Affordance grounding, a task to ground (i.e., localize) action possibility region in objects, which faces the challenge of establishing an explicit link with object parts due to the diversity of interactive affordance. Human has the ability…

Computer Vision and Pattern Recognition · Computer Science 2022-03-21 Hongchen Luo , Wei Zhai , Jing Zhang , Yang Cao , Dacheng Tao

This paper presents an unsupervised approach towards automatically extracting video-based guidance on object usage, from egocentric video and wearable gaze tracking, collected from multiple users while performing tasks. The approach i)…

Computer Vision and Pattern Recognition · Computer Science 2016-03-22 Dima Damen , Teesid Leelasawassuk , Walterio Mayol-Cuevas

Human activity recognition is typically addressed by detecting key concepts like global and local motion, features related to object classes present in the scene, as well as features related to the global context. The next open challenges…

Computer Vision and Pattern Recognition · Computer Science 2018-09-21 Fabien Baradel , Natalia Neverova , Christian Wolf , Julien Mille , Greg Mori

Active perception, the ability of a robot to proactively adjust its viewpoint to acquire task-relevant information, is essential for robust operation in unstructured real-world environments. While critical for downstream tasks such as…

Robotics · Computer Science 2026-03-03 Yongxi Huang , Zhuohang Wang , Wenjing Tang , Cewu Lu , Panpan Cai

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-Language-Action models…

Robotics · Computer Science 2025-09-29 Tomoya Yoshida , Shuhei Kurita , Taichi Nishimura , Shinsuke Mori

3D visual grounding aims to find the object within point clouds mentioned by free-form natural language descriptions with rich semantic cues. However, existing methods either extract the sentence-level features coupling all words or focus…

Computer Vision and Pattern Recognition · Computer Science 2023-10-09 Yanmin Wu , Xinhua Cheng , Renrui Zhang , Zesen Cheng , Jian Zhang

Today, mobile robots are expected to carry out increasingly complex tasks in multifarious, real-world environments. Often, the tasks require a certain semantic understanding of the workspace. Consider, for example, spoken instructions from…

Robotics · Computer Science 2014-01-21 Javier Velez , Garrett Hemann , Albert S. Huang , Ingmar Posner , Nicholas Roy

Automatic generation of textual video descriptions that are time-aligned with video content is a long-standing goal in computer vision. The task is challenging due to the difficulty of bridging the semantic gap between the visual and…

Computer Vision and Pattern Recognition · Computer Science 2018-09-25 Meera Hahn , Nataniel Ruiz , Jean-Baptiste Alayrac , Ivan Laptev , James M. Rehg

Object-goal navigation (Object-nav) entails searching, recognizing and navigating to a target object. Object-nav has been extensively studied by the Embodied-AI community, but most solutions are often restricted to considering static…

Recent work has presented embodied agents that can navigate to point-goal targets in novel indoor environments with near-perfect accuracy. However, these agents are equipped with idealized sensors for localization and take deterministic…

Computer Vision and Pattern Recognition · Computer Science 2020-09-08 Samyak Datta , Oleksandr Maksymets , Judy Hoffman , Stefan Lee , Dhruv Batra , Devi Parikh
‹ Prev 1 4 5 6 7 8 10 Next ›