English
Related papers

Related papers: Emergent Extreme-View Geometry in 3D Foundation Mo…

200 papers

Depth estimation is a foundational component for 3D reconstruction in minimally invasive endoscopic surgeries. However, existing monocular depth estimation techniques often exhibit limited performance to the varying illumination and complex…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Xinning Yao , Bo Liu , Bojian Li , Jingjing Wang , Jinghua Yue , Fugen Zhou

Vision-Language Models (VLMs) excel at 2D tasks such as grounding and captioning, yet remain limited in 3D understanding. A key limitation is their text-only supervision paradigm, which under-constrains fine-grained visual perception and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Hanxun Yu , Xuan Qu , Yuxin Wang , Jianke Zhu , Lei Ke

We introduce the first learning-based dense matching algorithm, termed Equirectangular Projection-Oriented Dense Kernelized Feature Matching (EDM), specifically designed for omnidirectional images. Equirectangular projection (ERP) images,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Dongki Jung , Jaehoon Choi , Yonghan Lee , Somi Jeong , Taejae Lee , Dinesh Manocha , Suyong Yeon

Depth and ego-motion estimations are essential for the localization and navigation of autonomous robots and autonomous driving. Recent studies make it possible to learn the per-pixel depth and ego-motion from the unlabeled monocular video.…

Computer Vision and Pattern Recognition · Computer Science 2022-06-09 Guangming Wang , Jiquan Zhong , Shijie Zhao , Wenhua Wu , Zhe Liu , Hesheng Wang

Human perception of 3D shapes goes beyond reconstructing them as a set of points or a composition of geometric primitives: we also effortlessly understand higher-level shape structure such as the repetition and reflective symmetry of object…

Computer Vision and Pattern Recognition · Computer Science 2019-08-13 Yonglong Tian , Andrew Luo , Xingyuan Sun , Kevin Ellis , William T. Freeman , Joshua B. Tenenbaum , Jiajun Wu

We present a deployment friendly, fast bottom-up framework for multi-person 3D human pose estimation. We adopt a novel neural representation of multi-person 3D pose which unifies the position of person instances with their corresponding 3D…

Computer Vision and Pattern Recognition · Computer Science 2020-08-05 Jogendra Nath Kundu , Ambareesh Revanur , Govind Vitthal Waghmare , Rahul Mysore Venkatesh , R. Venkatesh Babu

We address the 3D reconstruction of human faces from a single RGB image. To this end, we propose Pixel3DMM, a set of highly-generalized vision transformers which predict per-pixel geometric cues in order to constrain the optimization of a…

Computer Vision and Pattern Recognition · Computer Science 2025-05-02 Simon Giebenhain , Tobias Kirschstein , Martin Rünz , Lourdes Agapito , Matthias Nießner

We present FoundationPose, a unified foundation model for 6D object pose estimation and tracking, supporting both model-based and model-free setups. Our approach can be instantly applied at test-time to a novel object without fine-tuning,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-28 Bowen Wen , Wei Yang , Jan Kautz , Stan Birchfield

Learning model-free object pose estimation for unseen instances remains a fundamental challenge in 3D vision. Existing methods typically fall into two disjoint paradigms: category-level approaches predict absolute poses in a canonical space…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Weihang Li , Lorenzo Garattoni , Fabien Despinoy , Nassir Navab , Benjamin Busam

Pose estimation from unordered images is fundamental for 3D reconstruction, robotics, and scientific imaging. Recent geometric foundation models, such as DUSt3R, enable end-to-end dense 3D reconstruction but remain underexplored in…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Jiakai Zhang , Shouchen Zhou , Haizhao Dai , Xinhang Liu , Peihao Wang , Zhiwen Fan , Yuan Pei , Jingyi Yu

Feed-forward foundation models for multi-view 3-dimensional (3D) reconstruction have been trained on large-scale datasets of perspective images; when tested on wide field-of-view images, e.g., from a fisheye camera, their performance…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Ruxiao Duan , Erin Hong , Dongxu Zhao , Eric Turner , Alex Wong , Yunwen Zhou

We study the problem of applying 3D Foundation Models (3DFMs) to dense Novel View Synthesis (NVS). Despite significant progress in Novel View Synthesis powered by NeRF and 3DGS, current approaches remain reliant on accurate 3D attributes…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Yang Liu , Chuanchen Luo , Zimo Tang , Junran Peng , Zhaoxiang Zhang

State-of-the-art face super-resolution methods employ deep convolutional neural networks to learn a mapping between low- and high- resolution facial patterns by exploring local appearance knowledge. However, most of these methods do not…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Xiaobin Hu , Wenqi Ren , John LaMaster , Xiaochun Cao , Xiaoming Li , Zechao Li , Bjoern Menze , Wei Liu

In this paper, a novel deep-learning based framework is proposed to infer 3D human poses from a single image. Specifically, a two-phase approach is developed. We firstly utilize a generator with two branches for the extraction of explicit…

Computer Vision and Pattern Recognition · Computer Science 2018-09-24 Kun Zhou , Jinmiao Cai , Yao Li , Yulong Shi , Xiaoguang Han , Nianjuan Jiang , Kui Jia , Jiangbo Lu

Deep learning affords enormous opportunities to augment the armamentarium of biomedical imaging, albeit its design and implementation have potential flaws. Fundamentally, most deep learning models are driven entirely by data without…

Computer Vision and Pattern Recognition · Computer Science 2021-05-26 Liyue Shen , Wei Zhao , Dante Capaldi , John Pauly , Lei Xing

Estimating the 6D pose of novel objects is a fundamental yet challenging problem in robotics, often relying on access to object CAD models. However, acquiring such models can be costly and impractical. Recent approaches aim to bypass this…

Robotics · Computer Science 2025-08-25 Zhaodong Jiang , Ashish Sinha , Tongtong Cao , Yuan Ren , Bingbing Liu , Binbin Xu

We present a self-supervised learning algorithm for 3D human pose estimation of a single person based on a multiple-view camera system and 2D body pose estimates for each view. To train our model, represented by a deep neural network, we…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 Arij Bouazizi , Julian Wiederer , Ulrich Kressel , Vasileios Belagiannis

Central to the application of many multi-view geometry algorithms is the extraction of matching points between multiple viewpoints, enabling classical tasks such as camera pose estimation and 3D reconstruction. Many approaches that…

Computer Vision and Pattern Recognition · Computer Science 2021-10-15 Alexander Mai , Allen Yang , Dominique E. Meyer

Feed-forward 3D reconstruction models such as DUSt3R, VGGT, and Depth Anything 3 (DA3) are transformer-based foundation models that infer camera geometry and dense scene structure in a single forward pass. Trained at scale in a supervised…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jelena Bratulić , Sudhanshu Mittal , Thomas Brox , Christian Rupprecht

We present an approach for detecting and estimating the 3D poses of objects in images that requires only an untextured CAD model and no training phase for new objects. Our approach combines Deep Learning and 3D geometry: It relies on an…

Computer Vision and Pattern Recognition · Computer Science 2020-10-09 Giorgia Pitteri , Aurélie Bugeau , Slobodan Ilic , Vincent Lepetit