English
Related papers

Related papers: Perceive-then-Plan: Layout-as-Policy for Monocular…

200 papers

The definition of factor space and a unified optimization based classification model were developed for linear programming. Intelligent behaviour appeared in a decision process can be treated as a point y, the dynamic state observed and…

Optimization and Control · Mathematics 2021-01-12 Jing He , Qi-Wei Kong , Ho-Chung Lui , Hai-Tao Liu , Yi-Mu Ji , Hai-Chang Yao , Mo-Zhengfu Liu

We present an approach for the planar surface reconstruction of a scene from images with limited overlap. This reconstruction task is challenging since it requires jointly reasoning about single image 3D reconstruction, correspondence…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Samir Agarwala , Linyi Jin , Chris Rockwell , David F. Fouhey

Enabling agents to understand and interact with complex 3D scenes is a fundamental challenge for embodied artificial intelligence systems. While Multimodal Large Language Models (MLLMs) have achieved significant progress in 2D image…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Haoyuan Li , Rui Liu , Hehe Fan , Yi Yang

There is some ambiguity in the 3D shape of an object when the number of observed views is small. Because of this ambiguity, although a 3D object reconstructor can be trained using a single view or a few views per object, reconstructed…

Computer Vision and Pattern Recognition · Computer Science 2019-04-01 Hiroharu Kato , Tatsuya Harada

Monocular 3D reconstruction of articulated object categories is challenging due to the lack of training data and the inherent ill-posedness of the problem. In this work we use video self-supervision, forcing the consistency of consecutive…

Computer Vision and Pattern Recognition · Computer Science 2021-04-28 Filippos Kokkinos , Iasonas Kokkinos

Autonomous agents embedded in a physical environment need the ability to recognize objects and their properties from sensory data. Such a perceptual ability is often implemented by supervised machine learning models, which are pre-trained…

Layout estimation and 3D object detection are two fundamental tasks in indoor scene understanding. When combined, they enable the creation of a compact yet semantically rich spatial representation of a scene. Existing approaches typically…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Anton Konushin , Nikita Drozdov , Bulat Gabdullin , Alexey Zakharov , Anna Vorontsova , Danila Rukhovich , Maksim Kolodiazhnyi

3D line mapping from multi-view RGB images provides a compact and structured visual representation of scenes. We study the problem from a physical and topological perspective: a 3D line most naturally emerges as the edge of a finite 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Zeran Ke , Bin Tan , Gui-Song Xia , Yujun Shen , Nan Xue

Recent approaches for predicting layouts from 360 panoramas produce excellent results. These approaches build on a common framework consisting of three steps: a pre-processing step based on edge-based alignment, prediction of layout…

Computer Vision and Pattern Recognition · Computer Science 2020-12-29 Chuhang Zou , Jheng-Wei Su , Chi-Han Peng , Alex Colburn , Qi Shan , Peter Wonka , Hung-Kuo Chu , Derek Hoiem

Despite great strides in language-guided manipulation, existing work has been constrained to table-top settings. Table-tops allow for perfect and consistent camera angles, properties are that do not hold in mobile manipulation. Task plans…

Robotics · Computer Science 2023-11-08 Priyam Parashar , Vidhi Jain , Xiaohan Zhang , Jay Vakil , Sam Powers , Yonatan Bisk , Chris Paxton

Existing monocular 3D detectors typically tame the pronounced nonlinear regression of 3D bounding box through decoupled prediction paradigm, which employs multiple branches to estimate geometric center, depth, dimensions, and rotation angle…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Yifan Wang , Yian Zhao , Fanqi Pu , Xiaochen Yang , Yang Tang , Xi Chen , Wenming Yang

The representation of geometry in real-time 3D perception systems continues to be a critical research issue. Dense maps capture complete surface shape and can be augmented with semantic labels, but their high dimensionality makes them…

Computer Vision and Pattern Recognition · Computer Science 2019-04-16 Michael Bloesch , Jan Czarnowski , Ronald Clark , Stefan Leutenegger , Andrew J. Davison

Systems which incrementally create 3D semantic maps from image sequences must store and update representations of both geometry and semantic entities. However, while there has been much work on the correct formulation for geometrical…

Computer Vision and Pattern Recognition · Computer Science 2019-03-19 Shuaifeng Zhi , Michael Bloesch , Stefan Leutenegger , Andrew J. Davison

The 3D world limits the human body pose and the human body pose conveys information about the surrounding objects. Indeed, from a single image of a person placed in an indoor scene, we as humans are adept at resolving ambiguities of the…

Computer Vision and Pattern Recognition · Computer Science 2021-04-19 Zhenzhen Weng , Serena Yeung

Inferring the 3D structure from a single image, particularly in occluded regions, remains a fundamental yet unsolved challenge in vision-centric autonomous driving. Existing unsupervised approaches typically train a neural radiance field…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Zizhan Guo , Yi Feng , Mengtan Zhang , Haoran Zhang , Wei Ye , Rui Fan

Urban planning refers to the efforts of designing land-use configurations. Effective urban planning can help to mitigate the operational and social vulnerability of a urban system, such as high tax, crimes, traffic congestion and accidents,…

Artificial Intelligence · Computer Science 2021-01-08 Dongjie Wang , Yanjie Fu , Pengyang Wang , Bo Huang , Chang-Tien Lu

Building models that can understand and reason about 3D scenes is difficult owing to the lack of data sources for 3D supervised training and large-scale training regimes. In this work we ask - How can the knowledge in a pre-trained language…

Computer Vision and Pattern Recognition · Computer Science 2024-08-14 Shivam Chandhok

Abstract--- Exploiting the spatial structure in scene images is a key research direction for scene recognition. Due to the large intra-class structural diversity, building and modeling flexible structural layout to adapt various image…

Computer Vision and Pattern Recognition · Computer Science 2020-06-24 Gongwei Chen , Xinhang Song , Haitao Zeng , Shuqiang Jiang

Many existing autonomous driving paradigms involve a multi-stage discrete pipeline of tasks. To better predict the control signals and enhance user safety, an end-to-end approach that benefits from joint spatial-temporal feature learning is…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Shengchao Hu , Li Chen , Penghao Wu , Hongyang Li , Junchi Yan , Dacheng Tao

Most recent approaches to monocular 3D human pose estimation rely on Deep Learning. They typically involve regressing from an image to either 3D joint coordinates directly or 2D joint locations from which 3D coordinates are inferred. Both…

Computer Vision and Pattern Recognition · Computer Science 2017-04-11 Bugra Tekin , Pablo Márquez-Neila , Mathieu Salzmann , Pascal Fua