English
Related papers

Related papers: Grounding 3D Scene Affordance From Egocentric Inte…

200 papers

Generating 3D scenes from human motion sequences supports numerous applications, including virtual reality and architectural design. However, previous auto-regression-based human-aware 3D scene generation methods have struggled to…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Xiaolin Hong , Hongwei Yi , Fazhi He , Qiong Cao

Methods that synthesize indoor 3D scenes from text prompts have wide-ranging applications in film production, interior design, video games, virtual reality, and synthetic data generation for training embodied agents. Existing approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Antonio Ruiz , Tao Wu , Andrew Melnik , Qing Cheng , Xuqin Wang , Lu Liu , Yongliang Wang , Yanfeng Zhang , Helge Ritter

We study the task of embodied visual active learning, where an agent is set to explore a 3d environment with the goal to acquire visual scene understanding by actively selecting views for which to request annotation. While accurate on some…

Computer Vision and Pattern Recognition · Computer Science 2020-12-18 David Nilsson , Aleksis Pirinen , Erik Gärtner , Cristian Sminchisescu

Vision-Language Models (VLMs) have shown great success as foundational models for downstream vision and natural language applications in a variety of domains. However, these models are limited to reasoning over objects and actions currently…

Robotics · Computer Science 2025-06-13 Zachary Chavis , Hyun Soo Park , Stephen J. Guy

3D semantic occupancy prediction has become a crucial perception task for comprehensive scene understanding in autonomous driving. While recent advances have explored 3D Gaussian splatting for occupancy modeling to substantially reduce…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Xiaoyang Yan , Muleilan Pei , Shaojie Shen

Robotic mapping systems typically approach building metric-semantic scene representations from the robot's own sensors and cameras. However, these "first person" maps inherit the robot's own limitations due to its embodiment or skillset,…

Robotics · Computer Science 2026-03-31 Alan Yu , Yun Chang , Christopher Xie , Luca Carlone

Vision-language navigation (VLN) requires an agent to traverse complex 3D environments based on natural language instructions, necessitating a thorough scene understanding. While existing works equip agents with various scene…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Jianzhe Gao , Rui Liu , Wenguan Wang

Affordance refers to the functional properties that an agent perceives and utilizes from its environment, and is key perceptual information required for robots to perform actions. This information is rich and multimodal in nature. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Yizhou Huang , Fan Yang , Guoliang Zhu , Gen Li , Hao Shi , Yukun Zuo , Wenrui Chen , Zhiyong Li , Kailun Yang

Accurate 3D understanding of human hands and objects during manipulation remains a significant challenge for egocentric computer vision. Existing hand-object interaction datasets are predominantly captured in controlled studio settings,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Patrick Rim , Kevin Harris , Braden Copple , Shangchen Han , Xu Xie , Ivan Shugurov , Sizhe An , He Wen , Alex Wong , Tomas Hodan , Kun He

Visual grounding aims to identify objects or regions in a scene based on natural language descriptions, essential for spatially aware perception in autonomous driving. However, existing visual grounding tasks typically depend on bounding…

Computer Vision and Pattern Recognition · Computer Science 2025-09-04 Zhan Shi , Song Wang , Junbo Chen , Jianke Zhu

When humans perform a task with an articulated object, they interact with the object only in a handful of ways, while the space of all possible interactions is nearly endless. This is because humans have prior knowledge about what…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Liquan Wang , Nikita Dvornik , Rafael Dubeau , Mayank Mittal , Animesh Garg

Affordance grounding requires identifying where and how an agent should interact in open-world scenes, where actionable regions are often small, occluded, reflective, and visually ambiguous. Recent systems therefore combine multiple skills…

Robotics · Computer Science 2026-05-11 Haojian Huang , Jiahao Shi , Yinchuan Li , Yingcong Chen

3D Semantic Scene Graph Prediction aims to detect objects and their semantic relationships in 3D scenes, and has emerged as a crucial technology for robotics and AR/VR applications. While previous research has addressed dataset limitations…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 KunHo Heo , GiHyun Kim , SuYeon Kim , MyeongAh Cho

3D affordance segmentation aims to link human instructions to touchable regions of 3D objects for embodied manipulations. Existing efforts typically adhere to single-object, single-affordance paradigms, where each affordance type or…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Chunlin Yu , Hanqing Wang , Ye Shi , Haoyang Luo , Sibei Yang , Jingyi Yu , Jingya Wang

Semantics has enabled 3D scene understanding and affordance-driven object interaction. However, robots operating in real-world environments face a critical limitation: they cannot anticipate how objects move. Long-horizon mobile…

Representing and understanding 3D environments in a structured manner is crucial for autonomous agents to navigate and reason about their surroundings. While traditional Simultaneous Localization and Mapping (SLAM) methods generate metric…

Robotics · Computer Science 2026-02-03 Albert Gassol Puigjaner , Angelos Zacharia , Kostas Alexis

While Open Set Semantic Mapping and 3D Semantic Scene Graphs (3DSSGs) are established paradigms in robotic perception, deploying them effectively to support high-level reasoning in large-scale, real-world environments remains a significant…

Robotics · Computer Science 2026-02-04 Martin Günther , Felix Igelbrink , Oscar Lima , Lennart Niecksch , Marian Renz , Martin Atzmueller

Vision-language-action (VLA) models have shown strong potential for generalist robot manipulation, yet they remain limited by insufficient spatial reasoning, particularly in determining where to interact in complex visual scenes. While…

Robotics · Computer Science 2026-05-26 Runze Wang , Yuqian Fu , Yu Li , Tao Lin , Tianwen Qian , Mohamed Elhoseiny , Bo Zhao , Yanwei Fu , Yu-Gang Jiang , Xiangyang Xue

Egocentric interactive world models are essential for augmented reality and embodied AI, where visual generation must respond to user input with low latency, geometric consistency, and long-term stability. We study egocentric interaction…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Yuxi Wang , Wenqi Ouyang , Tianyi Wei , Yi Dong , Zhiqi Shen , Xingang Pan

The ability to understand the ways to interact with objects from visual cues, a.k.a. visual affordance, is essential to vision-guided robotic research. This involves categorizing, segmenting and reasoning of visual affordance. Relevant…

Computer Vision and Pattern Recognition · Computer Science 2021-04-01 Shengheng Deng , Xun Xu , Chaozheng Wu , Ke Chen , Kui Jia