English
Related papers

Related papers: Perceive-then-Plan: Layout-as-Policy for Monocular…

200 papers

As part of human core knowledge, the representation of objects is the building block of mental representation that supports high-level concepts and symbolic reasoning. While humans develop the ability of perceiving objects situated in 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-03-07 John Day , Tushar Arora , Jirui Liu , Li Erran Li , Ming Bo Cai

Learning to reconstruct 3D shapes using 2D images is an active research topic, with benefits of not requiring expensive 3D data. However, most work in this direction requires multi-view images for each object instance as training…

Computer Vision and Pattern Recognition · Computer Science 2021-07-14 Bo Peng , Wei Wang , Jing Dong , Tieniu Tan

Recognizing and localizing objects in the 3D space is a crucial ability for an AI agent to perceive its surrounding environment. While significant progress has been achieved with expensive LiDAR point clouds, it poses a great challenge for…

Computer Vision and Pattern Recognition · Computer Science 2021-08-16 Li Wang , Li Zhang , Yi Zhu , Zhi Zhang , Tong He , Mu Li , Xiangyang Xue

Given a monocular colour image of a warehouse rack, we aim to predict the bird's-eye view layout for each shelf in the rack, which we term as multi-layer layout prediction. To this end, we present RackLay, a deep neural network for…

Computer Vision and Pattern Recognition · Computer Science 2021-10-29 Meher Shashwat Nigam , Avinash Prabhu , Anurag Sahu , Puru Gupta , Tanvi Karandikar , N. Sai Shankar , Ravi Kiran Sarvadevabhatla , K. Madhava Krishna

We propose a novel approach to jointly perform 3D shape retrieval and pose estimation from monocular images.In order to make the method robust to real-world image variations, e.g. complex textures and backgrounds, we learn an embedding…

Computer Vision and Pattern Recognition · Computer Science 2019-03-28 Kyaw Zaw Lin , Weipeng Xu , Qianru Sun , Christian Theobalt , Tat-Seng Chua

3D scene generation conditioned on text prompts has significantly progressed due to the development of 2D diffusion generation models. However, the textual description of 3D scenes is inherently inaccurate and lacks fine-grained control…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Minglin Chen , Longguang Wang , Sheng Ao , Ye Zhang , Kai Xu , Yulan Guo

One major goal of vision is to infer physical models of objects, surfaces, and their layout from sensors. In this paper, we aim to interpret indoor scenes from one RGBD image. Our representation encodes the layout of walls, which must…

Computer Vision and Pattern Recognition · Computer Science 2017-08-21 Ruiqi Guo , Chuhang Zou , Derek Hoiem

Performing single image holistic understanding and 3D reconstruction is a central task in computer vision. This paper presents an integrated system that performs dense scene labeling, object detection, instance segmentation, depth…

Computer Vision and Pattern Recognition · Computer Science 2021-12-01 Sainan Liu , Vincent Nguyen , Yuan Gao , Subarna Tripathi , Zhuowen Tu

This paper presents a semantic planar SLAM system that improves pose estimation and mapping using cues from an instance planar segmentation network. While the mainstream approaches are using RGB-D sensors, employing a monocular camera with…

Computer Vision and Pattern Recognition · Computer Science 2022-06-23 Fangwen Shu , Yaxu Xie , Jason Rambach , Alain Pagani , Didier Stricker

SLAM systems are mainly applied for robot navigation while research on feasibility for motion planning with SLAM for tasks like bin-picking, is scarce. Accurate 3D reconstruction of objects and environments is important for planning motion…

Computer Vision and Pattern Recognition · Computer Science 2018-03-07 Sergey Triputen , Atmaraaj Gopal , Thomas Weber , Christian Hofert , Kristiaan Schreve , Matthias Ratsch

Well-designed indoor scenes should prioritize how people can act within a space rather than merely what objects to place. However, existing 3D scene generation methods emphasize visual and semantic plausibility, while insufficiently…

Human-Computer Interaction · Computer Science 2026-03-04 Semin Jin , Donghyuk Kim , Jeongmin Ryu , Kyung Hoon Hyun

We propose an unsupervised method for parsing large 3D scans of real-world scenes with easily-interpretable shapes. This work aims to provide a practical tool for analyzing 3D scenes in the context of aerial surveying and mapping, without…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Romain Loiseau , Elliot Vincent , Mathieu Aubry , Loic Landrieu

We present UniPlane, a novel method that unifies plane detection and reconstruction from posed monocular videos. Unlike existing methods that detect planes from local observations and associate them across the video for the final…

Computer Vision and Pattern Recognition · Computer Science 2024-07-08 Yuzhong Huang , Chen Liu , Ji Hou , Ke Huo , Shiyu Dong , Fred Morstatter

Recent advances in world models have shown promise for modeling future dynamics of environmental states, enabling agents to reason and act without accessing real environments. Current methods mainly perform single-step or fixed-horizon…

Computation and Language · Computer Science 2026-03-17 Youwei Liu , Jian Wang , Hanlin Wang , Beichen Guo , Wenjie Li

We present a damage-aware planning approach which determines the best sequence to manipulate a number of objects in a scene. This works on task-planning level, abstracts from motion planning and anticipates the dynamics of the scene using a…

Robotics · Computer Science 2016-12-05 Tobias Fromm , Andreas Birk

We present a new approach to the problem of estimating the 3D room layout from a single panoramic image. We represent room layout as three 1D vectors that encode, at each image column, the boundary positions of floor-wall and ceiling-wall,…

Computer Vision and Pattern Recognition · Computer Science 2019-04-09 Cheng Sun , Chi-Wei Hsiao , Min Sun , Hwann-Tzong Chen

A key challenge in the task of human pose and shape estimation is occlusion, including self-occlusions, object-human occlusions, and inter-person occlusions. The lack of diverse and accurate pose and shape training data becomes a major…

Computer Vision and Pattern Recognition · Computer Science 2022-03-02 Kaibing Yang , Renshu Gu , Maoyu Wang , Masahiro Toyoura , Gang Xu

Monocular dynamic reconstruction is a challenging and long-standing vision problem due to the highly ill-posed nature of the task. Existing approaches depend on templates, are effective only in quasi-static scenes, or fail to model 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Qianqian Wang , Vickie Ye , Hang Gao , Weijia Zeng , Jake Austin , Zhengqi Li , Angjoo Kanazawa

A major element of depth perception and 3D understanding is the ability to predict the 3D layout of a scene and its contained objects for a novel pose. Indoor environments are particularly suitable for novel view prediction, since the set…

Computer Vision and Pattern Recognition · Computer Science 2018-08-13 Pulak Purkait , Ujwal Bonde , Christopher Zach

We present a reward-predictive, model-based deep learning method featuring trajectory-constrained visual attention for local planning in visual navigation tasks. Our method learns to place visual attention at locations in latent image space…

Robotics · Computer Science 2022-05-27 Stefan Wapnick , Travis Manderson , David Meger , Gregory Dudek
‹ Prev 1 8 9 10 Next ›