中文
相关论文

相关论文: Grounding 3D Scene Affordance From Egocentric Inte…

200 篇论文

Can conversational videos captured from multiple egocentric viewpoints reveal the map of a scene in a cost-efficient way? We seek to answer this question by proposing a new problem: efficiently building the map of a previously unseen 3D…

计算机视觉与模式识别 · 计算机科学 2023-04-24 Sagnik Majumder , Hao Jiang , Pierre Moulon , Ethan Henderson , Paul Calamia , Kristen Grauman , Vamsi Krishna Ithapu

Incorporating the physical environment is essential for a complete understanding of human behavior in unconstrained every-day tasks. This is especially important in ego-centric tasks where obtaining 3 dimensional information is both…

计算机视觉与模式识别 · 计算机科学 2018-07-30 Mickey Li , Noyan Songur , Pavel Orlov , Stefan Leutenegger , A Aldo Faisal

This work tackles scene understanding for outdoor robotic navigation, solely relying on images captured by an on-board camera. Conventional visual scene understanding interprets the environment based on specific descriptive categories.…

机器人学 · 计算机科学 2022-02-07 Galadrielle Humblot-Renaux , Letizia Marchegiani , Thomas B. Moeslund , Rikke Gade

This paper addresses the problem of affordance grounding from RGBD images of an object, which aims to localize surface regions corresponding to a text query that describes an action on the object. While existing methods predict affordance…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Chunghyun Park , Seunghyeon Lee , Minsu Cho

If a robotic agent wants to exploit symbolic planning techniques to achieve some goal, it must be able to properly ground an abstract planning domain in the environment in which it operates. However, if the environment is initially unknown…

人工智能 · 计算机科学 2022-04-11 Leonardo Lamanna , Luciano Serafini , Alessandro Saetti , Alfonso Gerevini , Paolo Traverso

Egocentric videos present unique challenges for 3D scene understanding due to rapid camera motion, frequent object occlusions, and limited object visibility. This paper introduces a novel approach to instance segmentation and tracking in…

计算机视觉与模式识别 · 计算机科学 2024-11-21 Yash Bhalgat , Vadim Tschernezki , Iro Laina , João F. Henriques , Andrea Vedaldi , Andrew Zisserman

We introduce a novel task of reconstructing a time series of second-person 3D human body meshes from monocular egocentric videos. The unique viewpoint and rapid embodied camera motion of egocentric videos raise additional technical barriers…

计算机视觉与模式识别 · 计算机科学 2021-10-19 Miao Liu , Dexin Yang , Yan Zhang , Zhaopeng Cui , James M. Rehg , Siyu Tang

We address the problem of affordance reasoning in diverse scenes that appear in the real world. Affordances relate the agent's actions to their effects when taken on the surrounding objects. In our work, we take the egocentric view of the…

计算机视觉与模式识别 · 计算机科学 2018-06-18 Ching-Yao Chuang , Jiaman Li , Antonio Torralba , Sanja Fidler

Despite significant advancements in text-to-motion synthesis, generating language-guided human motion within 3D environments poses substantial challenges. These challenges stem primarily from (i) the absence of powerful generative models…

计算机视觉与模式识别 · 计算机科学 2024-03-28 Zan Wang , Yixin Chen , Baoxiong Jia , Puhao Li , Jinlu Zhang , Jingze Zhang , Tengyu Liu , Yixin Zhu , Wei Liang , Siyuan Huang

This article studies the commonsense object affordance concept for enabling close-to-human task planning and task optimization of embodied robotic agents in urban environments. The focus of the object affordance is on reasoning how to…

High-level autonomous operations depend on a robot's ability to construct a sufficiently expressive model of its environment. Traditional three-dimensional (3D) scene representations, such as point clouds and occupancy grids, provide…

机器人学 · 计算机科学 2025-06-10 Chad R Samuelson , Timothy W McLain , Joshua G Mangelson

Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object queries. In applications such as AI assistants, the perspective shifts -- inputs are…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Pengzhan Sun , Junbin Xiao , Tze Ho Elden Tse , Yicong Li , Arjun Akula , Angela Yao

Recent approaches have successfully focused on the segmentation of static reconstructions, thereby equipping downstream applications with semantic 3D understanding. However, the world in which we live is dynamic, characterized by numerous…

机器人学 · 计算机科学 2025-03-12 Tjark Behrens , René Zurbrügg , Marc Pollefeys , Zuria Bauer , Hermann Blum

Understanding how people interact with their surroundings and each other is essential for enabling robots to act in socially compliant and context-aware ways. While 3D Scene Graphs have emerged as a powerful semantic representation for…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Ermanno Bartoli , Dennis Rotondi , Buwei He , Patric Jensfelt , Kai O. Arras , Iolanda Leite

Egocentric scenes exhibit frequent occlusions, varied viewpoints, and dynamic interactions compared to typical scene understanding tasks. Occlusions and varied viewpoints can lead to multi-view semantic inconsistencies, while dynamic…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Di Li , Jie Feng , Jiahao Chen , Weisheng Dong , Guanbin Li , Guangming Shi , Licheng Jiao

A robot's ability to anticipate the 3D action target location of a hand's movement from egocentric videos can greatly improve safety and efficiency in human-robot interaction (HRI). While previous research predominantly focused on semantic…

机器人学 · 计算机科学 2024-03-11 Irving Fang , Yuzhong Chen , Yifan Wang , Jianghan Zhang , Qiushi Zhang , Jiali Xu , Xibo He , Weibo Gao , Hao Su , Yiming Li , Chen Feng

Perceiving and interacting with 3D articulated objects, such as cabinets, doors, and faucets, pose particular challenges for future home-assistant robots performing daily tasks in human environments. Besides parsing the articulated parts…

计算机视觉与模式识别 · 计算机科学 2023-05-05 Yian Wang , Ruihai Wu , Kaichun Mo , Jiaqi Ke , Qingnan Fan , Leonidas Guibas , Hao Dong

Egocentric sensors such as AR/VR devices capture human-object interactions and offer the potential to provide task-assistance by recalling 3D locations of objects of interest in the surrounding environment. This capability requires instance…

计算机视觉与模式识别 · 计算机科学 2024-06-10 Yunhan Zhao , Haoyu Ma , Shu Kong , Charless Fowlkes

Outdoor intelligent autonomous robotic operation relies on a sufficiently expressive map of the environment. Classical geometric mapping methods retain essential structural environment information, but lack a semantic understanding and…

We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently updates the underlying 3D scene state for continuous simulation. Existing egocentric simulators either…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Jinkun Hao , Mingda Jia , Ruiyan Wang , Xihui Liu , Ran Yi , Lizhuang Ma , Jiangmiao Pang , Xudong Xu