English
Related papers

Related papers: Towards Learning a Generalizable 3D Scene Represen…

200 papers

In the field of Robot Learning, the complex mapping between high-dimensional observations such as RGB images and low-level robotic actions, two inherently very different spaces, constitutes a complex learning problem, especially with…

Robotics · Computer Science 2024-05-29 Vitalis Vosylius , Younggyo Seo , Jafar Uruç , Stephen James

A robot operating in a household environment will see a wide range of unique and unfamiliar objects. While a system could train on many of these, it is infeasible to predict all the objects a robot will see. In this paper, we present a…

Robotics · Computer Science 2023-03-08 Ethan Chun , Yilun Du , Anthony Simeonov , Tomas Lozano-Perez , Leslie Kaelbling

Neural radiance fields (NeRFs) have emerged as a prominent pre-training paradigm for vision-centric autonomous driving, which enhances 3D geometry and appearance understanding in a fully self-supervised manner. To apply NeRF-based…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Hyeonjun Jeong , Juyeb Shin , Dongsuk Kum

Implicit representations such as Neural Radiance Fields (NeRF) have been shown to be very effective at novel view synthesis. However, these models typically require manual and careful human data collection for training. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2023-12-25 Pierre Marza , Laetitia Matignon , Olivier Simonin , Dhruv Batra , Christian Wolf , Devendra Singh Chaplot

As part of human core knowledge, the representation of objects is the building block of mental representation that supports high-level concepts and symbolic reasoning. While humans develop the ability of perceiving objects situated in 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-03-07 John Day , Tushar Arora , Jirui Liu , Li Erran Li , Ming Bo Cai

Generalizable perception is one of the pillars of high-level autonomy in space robotics. Estimating the structure and motion of unknown objects in dynamic environments is fundamental for such autonomous systems. Traditionally, the solutions…

Robotics · Computer Science 2024-11-26 Kuldeep R Barad , Antoine Richard , Jan Dentler , Miguel Olivares-Mendez , Carol Martinez

Global visual localization estimates the absolute pose of a camera using a single image, in a previously mapped area. Obtaining the pose from a single image enables many robotics and augmented/virtual reality applications. Inspired by…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Mohammad Altillawi , Shile Li , Sai Manoj Prakhya , Ziyuan Liu , Joan Serrat

We, as human beings, can understand and picture a familiar scene from arbitrary viewpoints given a single image, whereas this is still a grand challenge for computers. We hereby present a novel solution to mimic such human perception…

Computer Vision and Pattern Recognition · Computer Science 2022-05-06 Bangbang Yang , Yinda Zhang , Yijin Li , Zhaopeng Cui , Sean Fanello , Hujun Bao , Guofeng Zhang

In this study, we address the challenge of 3D scene structure recovery from monocular depth estimation. While traditional depth estimation methods leverage labeled datasets to directly predict absolute depth, recent advancements advocate…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Chi Zhang , Wei Yin , Gang Yu , Zhibin Wang , Tao Chen , Bin Fu , Joey Tianyi Zhou , Chunhua Shen

Most deep learning approaches to comprehensive semantic modeling of 3D indoor spaces require costly dense annotations in the 3D domain. In this work, we explore a central 3D scene modeling task, namely, semantic scene reconstruction without…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Junwen Huang , Alexey Artemov , Yujin Chen , Shuaifeng Zhi , Kai Xu , Matthias Nießner

Weakly-supervised 3D occupancy perception is crucial for vision-based autonomous driving in outdoor environments. Previous methods based on NeRF often face a challenge in balancing the number of samples used. Too many samples can decrease…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Qianpu Sun , Changyong Shu , Sifan Zhou , Runxi Cheng , Yongxian Wei , Zichen Yu , Dawei Yang , Sirui Han , Yuan Chun

Achieving generalizable and precise robotic manipulation across diverse environments remains a critical challenge, largely due to limitations in spatial perception. While prior imitation-learning approaches have made progress, their…

Robotics · Computer Science 2025-05-28 Yiqi Huang , Travis Davies , Jiahuan Yan , Jiankai Sun , Xiang Chen , Luhui Hu

In this paper, we tackle the challenge of predicting the unseen walls of a partially observed environment as a set of 2D line segments, conditioned on occupancy grids integrated along the trajectory of a 360{\deg} LIDAR sensor. A dataset of…

Robotics · Computer Science 2024-06-14 Ludvig Ericson , Patric Jensfelt

Robots rely heavily on sensors, especially RGB and depth cameras, to perceive and interact with the world. RGB cameras record 2D images with rich semantic information while missing precise spatial information. On the other side, depth…

Robotics · Computer Science 2023-10-16 Tong Zhang , Yingdong Hu , Hanchen Cui , Hang Zhao , Yang Gao

Predictive models can be particularly helpful for robots to effectively manipulate terrains in construction sites and extraterrestrial surfaces. However, terrain state representations become extremely high-dimensional especially to capture…

Robotics · Computer Science 2026-02-12 Chaoqi Liu , Yunzhu Li , Kris Hauser

Latent scene representation plays a significant role in training reinforcement learning (RL) agents. To obtain good latent vectors describing the scenes, recent works incorporate the 3D-aware latent-conditioned NeRF pipeline into scene…

Robotics · Computer Science 2024-09-30 Jiaxu Wang , Ziyi Zhang , Qiang Zhang , Jia Li , Jingkai Sun , Mingyuan Sun , Junhao He , Renjing Xu

Accurate scene perception is critical for vision-based robotic manipulation. Existing approaches typically follow either a Vision-to-Action (V-A) paradigm, predicting actions directly from visual inputs, or a Vision-to-3D-to-Action (V-3D-A)…

Robotics · Computer Science 2026-05-25 Ying Chai , Litao Deng , Ruizhi Shao , Jiajun Zhang , Kangchen Lv , Liangjun Xing , Xiang Li , Hongwen Zhang , Yebin Liu

We present a method to learn compositional multi-object dynamics models from image observations based on implicit object encoders, Neural Radiance Fields (NeRFs), and graph neural networks. NeRFs have become a popular choice for…

Computer Vision and Pattern Recognition · Computer Science 2022-07-28 Danny Driess , Zhiao Huang , Yunzhu Li , Russ Tedrake , Marc Toussaint

We present Im2Pano3D, a convolutional neural network that generates a dense prediction of 3D structure and a probability distribution of semantic labels for a full 360 panoramic view of an indoor scene when given only a partial observation…

Computer Vision and Pattern Recognition · Computer Science 2017-12-14 Shuran Song , Andy Zeng , Angel X. Chang , Manolis Savva , Silvio Savarese , Thomas Funkhouser

Location modeling, or determining where non-existing objects could feasibly appear in a scene, has the potential to benefit numerous computer vision tasks, from automatic object insertion to scene creation in virtual reality. Yet, this…

Computer Vision and Pattern Recognition · Computer Science 2025-10-08 Jooyeol Yun , Davide Abati , Mohamed Omran , Jaegul Choo , Amirhossein Habibian , Auke Wiggers