English
Related papers

Related papers: PanoVGGT: Feed-Forward 3D Reconstruction from Pano…

200 papers

Feedforward 3D Gaussian Splatting (3DGS) often struggles in trajectory-based sparse-view driving scenes. Existing Gaussian repair methods mainly target optimization-based 3DGS, while diffusion-based repair is typically restricted to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Rui Song , Tianhui Cai , Markus Gross , Xingcheng Zhou , Zewei Zhou , Zhiyu Huang , Olaf Wysocki , Jiaqi Ma

Generating ground-level views and coherent 3D site models from aerial-only imagery is challenging due to extreme viewpoint changes, missing intermediate observations, and large scale variations. Existing methods either refine renderings…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Sirshapan Mitra , Yogesh S. Rawat

Learning-based 3D reconstruction models, represented by Visual Geometry Grounded Transformers (VGGTs), have made remarkable progress with the use of large-scale transformers. Their prohibitive computational and memory costs severely hinder…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Weilun Feng , Haotong Qin , Mingqiang Wu , Chuanguang Yang , Yuqi Li , Xiangqi Li , Zhulin An , Libo Huang , Yulun Zhang , Michele Magno , Yongjun Xu

We propose SparseFusion, a sparse view 3D reconstruction approach that unifies recent advances in neural rendering and probabilistic image generation. Existing approaches typically build on neural rendering with re-projected features but…

Computer Vision and Pattern Recognition · Computer Science 2023-02-17 Zhizhuo Zhou , Shubham Tulsiani

3D room layout estimation by a single panorama using deep neural networks has made great progress. However, previous approaches can not obtain efficient geometry awareness of room layout with the only latitude of boundaries or…

Computer Vision and Pattern Recognition · Computer Science 2022-03-28 Zhigang Jiang , Zhongzheng Xiang , Jinhua Xu , Ming Zhao

Panoramic semantic segmentation models are typically trained under a strict gravity-aligned assumption. However, real-world captures often deviate from this canonical orientation due to unconstrained camera motions, such as the rotational…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Qinfeng Zhu , Yunxi Jiang , Lei Fan

In the field of novel-view synthesis, the necessity of knowing camera poses (e.g., via Structure from Motion) before rendering has been a common practice. However, the consistent acquisition of accurate camera poses remains elusive, and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Zhiwen Fan , Panwang Pan , Peihao Wang , Yifan Jiang , Hanwen Jiang , Dejia Xu , Zehao Zhu , Dilin Wang , Zhangyang Wang

3D reconstruction plays an increasingly important role in modern photogrammetric systems. Conventional satellite or aerial-based remote sensing (RS) platforms can provide the necessary data sources for the 3D reconstruction of large-scale…

Computer Vision and Pattern Recognition · Computer Science 2023-06-27 San Jiang , Kan You , Yaxin Li , Duojie Weng , Wu Chen

Panorama images have a much larger field-of-view thus naturally encode enriched scene context information compared to standard perspective images, which however is not well exploited in the previous scene understanding methods. In this…

Computer Vision and Pattern Recognition · Computer Science 2021-08-25 Cheng Zhang , Zhaopeng Cui , Cai Chen , Shuaicheng Liu , Bing Zeng , Hujun Bao , Yinda Zhang

We present PreF3R, Pose-Free Feed-forward 3D Reconstruction from an image sequence of variable length. Unlike previous approaches, PreF3R removes the need for camera calibration and reconstructs the 3D Gaussian field within a canonical…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Zequn Chen , Jiezhi Yang , Heng Yang

3D Gaussian Splatting is a powerful visual representation, providing high-quality and efficient 3D scene reconstruction, but it is crucially dependent on accurate camera poses typically obtained from computationally intensive processes like…

Robotics · Computer Science 2026-04-15 Daniel Yang , Jungseok Hong , John J. Leonard , Yogesh Girdhar

Recent approaches for predicting layouts from 360 panoramas produce excellent results. These approaches build on a common framework consisting of three steps: a pre-processing step based on edge-based alignment, prediction of layout…

Computer Vision and Pattern Recognition · Computer Science 2020-12-29 Chuhang Zou , Jheng-Wei Su , Chi-Han Peng , Alex Colburn , Qi Shan , Peter Wonka , Hung-Kuo Chu , Derek Hoiem

Reconstructing 3D scenes from sparse, unposed images remains challenging under real-world conditions with varying illumination and transient occlusions. Existing methods rely on scene-specific optimization using appearance embeddings or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Vinayak Gupta , Chih-Hao Lin , Shenlong Wang , Anand Bhattad , Jia-Bin Huang

Learning depth from spherical panoramas is becoming a popular research topic because a panorama has a full field-of-view of the environment and provides a relatively complete description of a scene. However, applying well-studied CNNs for…

Computer Vision and Pattern Recognition · Computer Science 2021-05-28 Hualie Jiang , Zhe Sheng , Siyu Zhu , Zilong Dong , Rui Huang

Recently, multiple formulations of vision problems as probabilistic inversions of generative models based on computer graphics have been proposed. However, applications to 3D perception from natural images have focused on low-dimensional…

Computer Vision and Pattern Recognition · Computer Science 2014-07-08 Tejas D. Kulkarni , Vikash K. Mansinghka , Pushmeet Kohli , Joshua B. Tenenbaum

The Visual Geometry Grounded Transformer (VGGT) enables strong feed-forward 3D reconstruction without per-scene optimization. However, its billion-parameter scale creates high memory and compute demands, hindering on-device deployment.…

Hardware Architecture · Computer Science 2026-01-29 Yipu Zhang , Jintao Cheng , Xingyu Liu , Zeyu Li , Carol Jingyi Li , Jin Wu , Lin Jiang , Yuan Xie , Jiang Xu , Wei Zhang

360{\deg} videos have emerged as a promising medium to represent our dynamic visual world. Compared to the "tunnel vision" of standard cameras, their borderless field of view offers a more complete perspective of our surroundings. While…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Rundong Luo , Matthew Wallingford , Ali Farhadi , Noah Snavely , Wei-Chiu Ma

Wide-baseline panoramic images are frequently used in applications like VR and simulations to minimize capturing labor costs and storage needs. However, synthesizing novel views from these panoramic images in real time remains a significant…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Zheng Chen , Chenming Wu , Zhelun Shen , Chen Zhao , Weicai Ye , Haocheng Feng , Errui Ding , Song-Hai Zhang

Reconstruction of 3D neural fields from posed images has emerged as a promising method for self-supervised representation learning. The key challenge preventing the deployment of these 3D scene learners on large-scale video data is their…

Computer Vision and Pattern Recognition · Computer Science 2023-06-02 Cameron Smith , Yilun Du , Ayush Tewari , Vincent Sitzmann

We introduce the Visual Implicit Geometry Transformer (ViGT), an autonomous driving geometric model that estimates continuous 3D occupancy fields from surround-view camera rigs. ViGT represents a step towards foundational geometric models…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Arsenii Shirokov , Mikhail Kuznetsov , Danila Stepochkin , Egor Evdokimov , Daniil Glazkov , Nikolay Patakin , Anton Konushin , Dmitry Senushkin