English
Related papers

Related papers: AREA3D: Active Reconstruction Agent with Unified F…

200 papers

Reconstructing 3D representations from 2D inputs is a fundamental task in computer vision and graphics, serving as a cornerstone for understanding and interacting with the physical world. While traditional methods achieve high fidelity,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Weijie Wang , Qihang Cao , Sensen Gao , Donny Y. Chen , Haofei Xu , Wenjing Bian , Songyou Peng , Tat-Jen Cham , Chuanxia Zheng , Andreas Geiger , Jianfei Cai , Jia-Wang Bian , Bohan Zhuang

Recent advances in 2D-to-3D perception have enabled the recovery of 3D scene semantics from unposed images. However, prevailing methods often suffer from limited generalization, reliance on per-scene optimization, and semantic…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Jie Hu , Shizun Wang , Xinchao Wang

Active vision is inherently attention-driven: The agent actively selects views to attend in order to fast achieve the vision task while improving its internal representation of the scene being observed. Inspired by the recent success of…

Computer Vision and Pattern Recognition · Computer Science 2022-01-12 Min Liu , Yifei Shi , Lintao Zheng , Kai Xu , Hui Huang , Dinesh Manocha

In this paper, we rethink the problem of scene reconstruction from an embodied agent's perspective: While the classic view focuses on the reconstruction accuracy, our new perspective emphasizes the underlying functions and constraints such…

Robotics · Computer Science 2021-03-31 Muzhi Han , Zeyu Zhang , Ziyuan Jiao , Xu Xie , Yixin Zhu , Song-Chun Zhu , Hangxin Liu

We present Spatial Region 3D (SR-3D) aware vision-language model that connects single-view 2D images and multi-view 3D data through a shared visual token space. SR-3D supports flexible region prompting, allowing users to annotate regions…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 An-Chieh Cheng , Yang Fu , Yukang Chen , Zhijian Liu , Xiaolong Li , Subhashree Radhakrishnan , Song Han , Yao Lu , Jan Kautz , Pavlo Molchanov , Hongxu Yin , Xiaolong Wang , Sifei Liu

The ability to accurately reconstruct the 3D facets of a scene is one of the key problems in robotic vision. However, even with recent advances with machine learning, there is no high-fidelity universal 3D reconstruction method for this…

Computer Vision and Pattern Recognition · Computer Science 2019-10-08 Bipul Islam , Ji Liu , Anthony Yezzi , Romeil Sandhu

We introduce a new method that efficiently computes a set of viewpoints and trajectories for high-quality 3D reconstructions in outdoor environments. Our goal is to automatically explore an unknown area, and obtain a complete 3D scan of a…

Computer Vision and Pattern Recognition · Computer Science 2018-09-19 Benjamin Hepp , Matthias Nießner , Otmar Hilliges

The rapid development of Large Multimodal Models (LMMs) has led to remarkable progress in 2D visual understanding; however, extending these capabilities to 3D scene understanding remains a significant challenge. Existing approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Hongpei Zheng , Lintao Xiang , Qijun Yang , Qian Lin , Hujun Yin

Vision-language models (VLMs) have achieved strong performance in multimodal understanding and reasoning, yet grounded reasoning in 3D scenes remains underexplored. Effective 3D reasoning hinges on accurate grounding: to answer open-ended…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Henry Zheng , Chenyue Fang , Rui Huang , Siyuan Wei , Xiao Liu , Gao Huang

Articulated objects, such as laptops and drawers, exhibit significant challenges for 3D reconstruction and pose estimation due to their multi-part geometries and variable joint configurations, which introduce structural diversity across…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 WenBo Xu , Liu Liu , Li Zhang , Ran Zhang , Hao Wu , Dan Guo , Meng Wang

Spatial reasoning in large-scale 3D environments remains challenging for current vision-language models, which are typically constrained to room-scale scenarios. We introduce H$^2$U3D (Holistic House Understanding in 3D), a 3D visual…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Hongpei Zheng , Shijie Li , Yanran Li , Hujun Yin

The ability to understand and reason the 3D real world is a crucial milestone towards artificial general intelligence. The current common practice is to finetune Large Language Models (LLMs) with 3D data and texts to enable 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Sha Zhang , Di Huang , Jiajun Deng , Shixiang Tang , Wanli Ouyang , Tong He , Yanyong Zhang

Multi-view 3D reconstruction has achieved remarkable progress with the advent of feed-forward 3D reconstruction models. However, these models are typically trained and evaluated under ideal, degradation-free imaging conditions, whereas…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Jin Hyeon Kim , Jaeeun Lee , Claire Kim , Kyoungjin Oh , Paul Hyunbin Cho , Jaewon Min , Yeji Choi , Jihye Park , Hyunhee Park , Minkyu Park , Seungryong Kim

3D Anomaly Detection (AD) is a promising means of controlling the quality of manufactured products. However, existing methods typically require carefully training a task-specific model for each category independently, leading to high cost,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Jiayi Cheng , Can Gao , Jie Zhou , Jiajun Wen , Tao Dai , Jinbao Wang

Dense 3D reconstruction and ego-motion estimation are key challenges in autonomous driving and robotics. Compared to the complex, multi-modal systems deployed today, multi-camera systems provide a simpler, low-cost alternative. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Aron Schmied , Tobias Fischer , Martin Danelljan , Marc Pollefeys , Fisher Yu

Recent advancements in 3D robotic manipulation have improved grasping of everyday objects, but transparent and specular materials remain challenging due to depth sensing limitations. While several 3D reconstruction and depth completion…

Robotics · Computer Science 2025-06-23 Mingxu Zhang , Xiaoqi Li , Jiahui Xu , Kaichen Zhou , Hojin Bae , Yan Shen , Chuyan Xiong , Hao Dong

Modeling 3D articulated objects with realistic geometry, textures, and kinematics is essential for a wide range of applications. However, existing optimization-based reconstruction methods often require dense multi-view inputs and expensive…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Sylvia Yuan , Ruoxi Shi , Xinyue Wei , Xiaoshuai Zhang , Hao Su , Minghua Liu

Implicit neural representations have shown compelling results in offline 3D reconstruction and also recently demonstrated the potential for online SLAM systems. However, applying them to autonomous 3D reconstruction, where a robot is…

Computer Vision and Pattern Recognition · Computer Science 2023-02-09 Yunlong Ran , Jing Zeng , Shibo He , Lincheng Li , Yingfeng Chen , Gimhee Lee , Jiming Chen , Qi Ye

Mapping and understanding complex 3D environments is fundamental to how autonomous systems perceive and interact with the physical world, requiring both precise geometric reconstruction and rich semantic comprehension. While existing 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Naman Patel , Prashanth Krishnamurthy , Farshad Khorrami

3D geometry is a very informative cue when interacting with and navigating an environment. This writing proposes a new approach to 3D reconstruction and scene understanding, which implicitly learns 3D geometry from depth maps pairing a deep…

Computer Vision and Pattern Recognition · Computer Science 2018-08-22 Dario Rethage , Federico Tombari , Felix Achilles , Nassir Navab