中文
相关论文

相关论文: Perceive-then-Plan: Layout-as-Policy for Monocular…

200 篇论文

Holistic 3D human-scene reconstruction is a crucial and emerging research area in robot perception. A key challenge in holistic 3D human-scene reconstruction is to generate a physically plausible 3D scene from a single monocular RGB image.…

计算机视觉与模式识别 · 计算机科学 2023-07-28 Sandika Biswas , Kejie Li , Biplab Banerjee , Subhasis Chaudhuri , Hamid Rezatofighi

We present a new pipeline for holistic 3D scene understanding from a single image, which could predict object shapes, object poses, and scene layout. As it is a highly ill-posed problem, existing methods usually suffer from inaccurate…

计算机视觉与模式识别 · 计算机科学 2021-08-24 Cheng Zhang , Zhaopeng Cui , Yinda Zhang , Bing Zeng , Marc Pollefeys , Shuaicheng Liu

3D scene understanding has gained significant attention due to its wide range of applications. However, existing methods for 3D scene understanding are limited to specific downstream tasks, which hinders their practicality in real-world…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Zehan Wang , Haifeng Huang , Yang Zhao , Ziang Zhang , Zhou Zhao

Pretraining 3D encoders by aligning with Contrastive Language Image Pretraining (CLIP) has emerged as a promising direction to learn generalizable representations for 3D scene understanding. In this paper, we propose UniScene3D, a…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Ye Mao , Weixun Luo , Ranran Huang , Junpeng Jing , Krystian Mikolajczyk

3D vision-language (VL) reasoning has gained significant attention due to its potential to bridge the 3D physical world with natural language descriptions. Existing approaches typically follow task-specific, highly specialized paradigms.…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Hao Liu , Yanni Ma , Yan Liu , Haihong Xiao , Ying He

We consider a category-level perception problem, where one is given 2D or 3D sensor data picturing an object of a given category (e.g., a car), and has to reconstruct the 3D pose and shape of the object despite intra-class variability…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Jingnan Shi , Heng Yang , Luca Carlone

This paper introduces Scene-LLM, a 3D-visual-language model that enhances embodied agents' abilities in interactive 3D indoor environments by integrating the reasoning strengths of Large Language Models (LLMs). Scene-LLM adopts a hybrid 3D…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Rao Fu , Jingyu Liu , Xilun Chen , Yixin Nie , Wenhan Xiong

In this paper, we demonstrate that mobile manipulation policies utilizing a 3D latent map achieve stronger spatial and temporal reasoning than policies relying solely on images. We introduce Seeing the Bigger Picture (SBP), an end-to-end…

机器人学 · 计算机科学 2026-03-06 Sunghwan Kim , Woojeh Chung , Zhirui Dai , Dwait Bhatt , Arth Shukla , Hao Su , Yulun Tian , Nikolay Atanasov

3D layout tasks have traditionally concentrated on geometric constraints, but many practical applications demand richer contextual understanding that spans social interactions, cultural traditions, and usage conventions. Existing methods…

图形学 · 计算机科学 2025-04-01 Yuto Asano , Naruya Kondo , Tatsuki Fushimi , Yoichi Ochiai

Just as humans can become disoriented in featureless deserts or thick fogs, not all environments are conducive to the Localization Accuracy and Stability (LAS) of autonomous robots. This paper introduces an efficient framework designed to…

机器人学 · 计算机科学 2024-08-06 Kaixin Chai , Long Xu , Qianhao Wang , Chao Xu , Peng Yin , Fei Gao

Predicting 3D room layout from single image is a challenging task with many applications. In this paper, we propose a new training and post-processing method for 3D room layout estimation, built on a recent state-of-the-art 3D room layout…

计算机视觉与模式识别 · 计算机科学 2020-09-08 Dongho Choi

Monocular 3D lane detection is a fundamental task in autonomous driving. Although sparse-point methods lower computational load and maintain high accuracy in complex lane geometries, current methods fail to fully leverage the geometric…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Yifan Chang , Junjie Huang , Xiaofeng Wang , Yun Ye , Zhujin Liang , Yi Shan , Dalong Du , Xingang Wang

We present ONCE-3DLanes, a real-world autonomous driving dataset with lane layout annotation in 3D space. Conventional 2D lane detection from a monocular image yields poor performance of following planning and control tasks in autonomous…

计算机视觉与模式识别 · 计算机科学 2022-05-17 Fan Yan , Ming Nie , Xinyue Cai , Jianhua Han , Hang Xu , Zhen Yang , Chaoqiang Ye , Yanwei Fu , Michael Bi Mi , Li Zhang

We introduce a new approach for estimating the 3D pose and the 3D shape of an object from a single image. Given a training set of view exemplars, we learn and select appearance-based discriminative parts which are mapped onto the 3D model…

计算机视觉与模式识别 · 计算机科学 2015-02-03 Menglong Zhu , Xiaowei Zhou , Kostas Daniilidis

Editing a 3D indoor scene from natural language is conceptually straightforward but technically challenging. Existing open-vocabulary systems often regenerate large portions of a scene or rely on image-space edits that disrupt spatial…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Seongrae Noh , SeungWon Seo , Gyeong-Moon Park , HyeongYeop Kang

We introduce the first learning-based reconstructability predictor to improve view and path planning for large-scale 3D urban scene acquisition using unmanned drones. In contrast to previous heuristic approaches, our method learns a model…

图形学 · 计算机科学 2022-09-22 Yilin Liu , Liqiang Lin , Yue Hu , Ke Xie , Chi-Wing Fu , Hao Zhang , Hui Huang

Scene flow estimation has been receiving increasing attention for 3D environment perception. Monocular scene flow estimation -- obtaining 3D structure and 3D motion from two temporally consecutive images -- is a highly ill-posed problem,…

计算机视觉与模式识别 · 计算机科学 2020-04-17 Junhwa Hur , Stefan Roth

Motivated by the advances in 3D sensing technology and the spreading of low-cost robotic platforms, 3D object reconstruction has become a common task in many areas. Nevertheless, the selection of the optimal sensor pose that maximizes the…

计算机视觉与模式识别 · 计算机科学 2021-01-27 Miguel Mendoza , J. Irving Vasquez-Gomez , Hind Taud , Luis Enrique Sucar , Carolina Reta

With the advent of large-scale pre-trained models, interest in adapting and exploiting them for continual learning scenarios has grown. In this paper, we propose an approach to exploiting pre-trained vision-language models (e.g. CLIP) that…

计算机视觉与模式识别 · 计算机科学 2023-11-01 Xialei Liu , Xusheng Cao , Haori Lu , Jia-wen Xiao , Andrew D. Bagdanov , Ming-Ming Cheng

Existing generative approaches for guided image synthesis of multi-object scenes typically rely on 2D controls in the image or text space. As a result, these methods struggle to maintain and respect consistent three-dimensional geometric…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Léopold Maillard , Tom Durand , Adrien Ramanana Rahary , Maks Ovsjanikov