English
Related papers

Related papers: G3T Up! Gravity Aligned Coordinate Frames Simplify…

200 papers

Reconstructing physically plausible 3D human-scene interactions (HSI) from a single image currently presents a trade-off: optimization based methods offer accurate contact but are slow (~20s), while feed-forward approaches are fast yet lack…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Pradyumna YM , Yuxuan Xue , Yue Chen , Nikita Kister , István Sárándi , Gerard Pons-Moll

3D semantic occupancy prediction has become a crucial perception task for comprehensive scene understanding in autonomous driving. While recent advances have explored 3D Gaussian splatting for occupancy modeling to substantially reduce…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Xiaoyang Yan , Muleilan Pei , Shaojie Shen

Monocular 3D human pose estimation remains a challenging and ill-posed problem, particularly in real-time settings and unconstrained environments. While direct imageto-3D approaches require large annotated datasets and heavy models,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Mohamed Adjel

In this work, we address the task of 3D reconstruction in dynamic scenes, where object motions frequently degrade the quality of previous 3D pointmap regression methods, such as DUSt3R, that are originally designed for static 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Jisang Han , Honggyu An , Jaewoo Jung , Takuya Narihira , Junyoung Seo , Kazumi Fukuda , Chaehyun Kim , Sunghwan Hong , Yuki Mitsufuji , Seungryong Kim

Depth estimation plays a pivotal role in autonomous driving, facilitating a comprehensive understanding of the vehicle's 3D surroundings. Radar, with its robustness to adverse weather conditions and capability to measure distances, has…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Huawei Sun , Zixu Wang , Hao Feng , Julius Ott , Lorenzo Servadei , Robert Wille

We present the first approach to build hierarchical task-driven 3D scene graphs of arbitrary indoor or outdoor environments using an uncalibrated monocular camera in real-time. We leverage geometric foundation models to estimate geometric…

Robotics · Computer Science 2026-05-26 Dominic Maggio , Nicolas Gorlo , Luca Carlone

Multi-beam LiDAR sensors, as used on autonomous vehicles and mobile robots, acquire sequences of 3D range scans ("frames"). Each frame covers the scene sparsely, due to limited angular scanning resolution and occlusion. The sparsity…

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Shengyu Huang , Zan Gojcic , Jiahui Huang , Andreas Wieser , Konrad Schindler

The choice of data representation is a key factor in the success of deep learning in geometric tasks. For instance, DUSt3R recently introduced the concept of viewpoint-invariant point maps, generalizing depth prediction and showing that all…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Ben Kaye , Tomas Jakab , Shangzhe Wu , Christian Rupprecht , Andrea Vedaldi

Existing RGB-based imitation learning approaches typically employ traditional vision encoders such as ResNet or ViT, which lack explicit 3D reasoning capabilities. Recent geometry-grounded vision models, such as VGGT~\cite{wang2025vggt},…

Robotics · Computer Science 2025-09-22 An Dinh Vuong , Minh Nhat Vu , Ian Reid

Parameter-efficient fine-tuning strategies for foundation models in 1D textual and 2D visual analysis have demonstrated remarkable efficacy. However, due to the scarcity of point cloud data, pre-training large 3D models remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Mengke Li , Lihao Chen , Peng Zhang , Yiu-ming Cheung , Hui Huang

With advanced X-ray source and detector technologies being continuously developed, non-traditional CT geometries have been widely explored. Generalized-Equiangular Geometry CT (GEGCT) architecture, in which an X-ray source might be…

Medical Physics · Physics 2023-06-30 Yingxian Xia , Zhiqiang Chen , Li Zhang , Yuxiang Xing , Hewei Gao

Matching 3D rigid point clouds in complex environments robustly and accurately is still a core technique used in many applications. This paper proposes a new architecture combining error estimation from sample covariances and dual global…

Computer Vision and Pattern Recognition · Computer Science 2017-07-28 Can Pu , Nanbo Li , Robert B Fisher

Feed-forward 3D reconstruction has advanced rapidly, but current models remain unreliable in UAV photogrammetric acquisition. We argue that this failure is caused not only by appearance-domain shift, but also by UAV-specific camera-geometry…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Xiang Yang , Yongli Wang , HaiFeng Li , Yunsheng Zhang

Joint camera pose and dense geometry estimation from a set of images or a monocular video remains a challenging problem due to its computational complexity and inherent visual ambiguities. Most dense incremental reconstruction systems…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Kirill Mazur , Gwangbin Bae , Andrew J. Davison

Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have advanced 3D reconstruction and novel view synthesis, but remain heavily dependent on accurate camera poses and dense viewpoint coverage. These requirements limit their…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Jiahui Lu , Haihong Xiao , Xueyan Zhao , Wenxiong Kang

The integration of aerial and ground images has been a promising solution in 3D modeling of complex scenes, which is seriously restricted by finding reliable correspondences. The primary contribution of this study is a feature matching…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Jiangxue Yu , Hui Wang , San Jiang , Xing Zhang , Dejin Zhang , Qingquan Li

Reconstructing dynamic 4D scenes from monocular videos is a fundamental yet challenging task. While recent 3D foundation models provide strong geometric priors, their performance significantly degrades in dynamic environments. This…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Ying Zang , Xuanyi Liu , Yidong Han , Deyi Ji , Chaotao Ding , Yuanqi Hu , Qi Zhu , Xuanfu Li , Jin Ma , Lingyun Sun , Tianrun Chen , Lanyun Zhu

Feed-forward 3D reconstruction models such as DUSt3R, VGGT, and Depth Anything 3 (DA3) are transformer-based foundation models that infer camera geometry and dense scene structure in a single forward pass. Trained at scale in a supervised…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jelena Bratulić , Sudhanshu Mittal , Thomas Brox , Christian Rupprecht

Existing methods for image alignment struggle in cases involving feature-sparse regions, extreme scale and field-of-view differences, and large deformations, often resulting in suboptimal accuracy. Robustness to these challenges can be…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Kanggeon Lee , Soochahn Lee , Kyoung Mu Lee

Time varying sequences of 3D point clouds, or 4D point clouds, are now being acquired at an increasing pace in several applications (e.g., LiDAR in autonomous or assisted driving). In many cases, such volume of data is transmitted, thus…

Computer Vision and Pattern Recognition · Computer Science 2023-06-05 Lorenzo Berlincioni , Stefano Berretti , Marco Bertini , Alberto Del Bimbo