English
Related papers

Related papers: PanoVGGT: Feed-Forward 3D Reconstruction from Pano…

200 papers

Panoramic depth estimation provides a comprehensive solution for capturing complete $360^\circ$ environmental structural information, offering significant benefits for robotics and AR/VR applications. However, while extensively studied in…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Hualie Jiang , Ziyang Song , Zhiqiang Lou , Rui Xu , Minglang Tan

Vision-centric occupancy networks, which represent the surrounding environment with uniform voxels with semantics, have become a new trend for safe driving of camera-only autonomous driving perception systems, as they are able to detect…

Computer Vision and Pattern Recognition · Computer Science 2024-09-19 Yining Shi , Jiusi Li , Kun Jiang , Ke Wang , Yunlong Wang , Mengmeng Yang , Diange Yang

Reconstructing 3D objects from a single image is an intriguing but challenging problem. One promising solution is to utilize multi-view (MV) 3D reconstruction to fuse generated MV images into consistent 3D objects. However, the generated…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Yizheng Chen , Rengan Xie , Qi Ye , Sen Yang , Zixuan Xie , Tianxiao Chen , Rong Li , Yuchi Huo

Existing image foundation models are not optimized for spherical images having been trained primarily on perspective images. PanoSAMic integrates the pre-trained Segment Anything (SAM) encoder to make use of its extensive training and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Mahdi Chamseddine , Didier Stricker , Jason Rambach

Visual place recognition has gained significant attention in recent years as a crucial technology in autonomous driving and robotics. Currently, the two main approaches are the perspective view retrieval (P2P) paradigm and the…

Computer Vision and Pattern Recognition · Computer Science 2023-07-31 Ze Shi , Hao Shi , Kailun Yang , Zhe Yin , Yining Lin , Kaiwei Wang

We present a fast, spatio-temporal scene understanding framework based on Visual Geometry Grounded Transformer (VGGT). The proposed pipeline is designed to enable efficient, close to real-time performance, supporting applications including…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Gergely Dinya , Péter Halász , András Lőrincz , Kristóf Karacs , Anna Gelencsér-Horváth

Recovering high-fidelity 3D hand geometry from images is a critical task in computer vision, holding significant value for domains such as robotics, animation and VR/AR. Crucially, scalable applications demand both accuracy and deployment…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Yumeng Liu , Xiao-Xiao Long , Marc Habermann , Xuanze Yang , Cheng Lin , Yuan Liu , Yuexin Ma , Wenping Wang , Ligang Liu

Despite remarkable progress in image translation, the complex scene with multiple discrepant objects remains a challenging problem. The translated images have low fidelity and tiny objects in fewer details causing unsatisfactory performance…

Computer Vision and Pattern Recognition · Computer Science 2022-12-26 Liyun Zhang , Photchara Ratsamee , Bowen Wang , Zhaojie Luo , Yuki Uranishi , Manabu Higashida , Haruo Takemura

We propose Flash3D, a method for scene reconstruction and novel view synthesis from a single image which is both very generalisable and efficient. For generalisability, we start from a "foundation" model for monocular depth estimation and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Stanislaw Szymanowicz , Eldar Insafutdinov , Chuanxia Zheng , Dylan Campbell , João F. Henriques , Christian Rupprecht , Andrea Vedaldi

Visual navigation and three-dimensional (3D) scene reconstruction are essential for robotics to interact with the surrounding environment. Large-scale scenes and critical camera motions are great challenges facing the research community to…

Computer Vision and Pattern Recognition · Computer Science 2021-09-21 Qi Cai , Lilian Zhang , Yuanxin Wu , Wenxian Yu , Dewen Hu

Streaming 3D reconstruction aims to recover 3D information, such as camera poses and point clouds, from a video stream, which necessitates geometric accuracy, temporal consistency, and computational efficiency. Motivated by the principles…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Lin-Zhuo Chen , Jian Gao , Yihang Chen , Ka Leong Cheng , Yipengjing Sun , Liangxiao Hu , Nan Xue , Xing Zhu , Yujun Shen , Yao Yao , Yinghao Xu

In this work, we introduce panoramic panoptic segmentation, as the most holistic scene understanding, both in terms of Field of View (FoV) and image-level understanding for standard camera-based input. A complete surrounding understanding…

Computer Vision and Pattern Recognition · Computer Science 2022-12-29 Alexander Jaus , Kailun Yang , Rainer Stiefelhagen

In this work, we introduce panoramic panoptic segmentation as the most holistic scene understanding both in terms of field of view and image level understanding for standard camera based input. A complete surrounding understanding provides…

Computer Vision and Pattern Recognition · Computer Science 2021-05-31 Alexander Jaus , Kailun Yang , Rainer Stiefelhagen

Human perceive the 3D world through 2D observations from limited viewpoints. While recent feed-forward generalizable 3D reconstruction models excel at recovering 3D structures from sparse images, their representations are often confined to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Mochu Xiang , Zhelun Shen , Xuesong Li , Jiahui Ren , Jing Zhang , Chen Zhao , Shanshan Liu , Haocheng Feng , Jingdong Wang , Yuchao Dai

Panoptic segmentation of 3D scenes, involving the segmentation and classification of object instances in a dense 3D reconstruction of a scene, is a challenging problem, especially when relying solely on unposed 2D images. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Lojze Zust , Yohann Cabon , Juliette Marrie , Leonid Antsfeld , Boris Chidlovskii , Jerome Revaud , Gabriela Csurka

We present Gen3R, a method that bridges the strong priors of foundational reconstruction models and video diffusion models for scene-level 3D generation. We repurpose the VGGT reconstruction model to produce geometric latents by training an…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Jiaxin Huang , Yuanbo Yang , Bangbang Yang , Lin Ma , Yuewen Ma , Yiyi Liao

Event cameras offer promising advantages such as high dynamic range and low latency, making them well-suited for challenging lighting conditions and fast-moving scenarios. However, reconstructing 3D scenes from raw event streams is…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Jiaxu Wang , Junhao He , Ziyi Zhang , Mingyuan Sun , Jingkai Sun , Renjing Xu

Reconstructing large-scale urban scenes from sparse aerial views is a crucial yet challenging task. Due to biased top-down and shallow-oblique camera poses, sparse aerial captures exhibit strong evidence imbalance: roofs and open regions…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Dongli Wu , Zhuoxiao Li , Tongyan Hua , Yinrui Ren , Xiaobao Wei , Rongjun Qin , Wufan Zhao

Image editing has made great progress on planar images, but panoramic image editing remains underexplored. Due to their spherical geometry and projection distortions, panoramic images present three key challenges: boundary discontinuity,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Zhiao Feng , Xuewei Li , Junjie Yang , Jingchao Li , Yuxin Peng , Xi Li

Convolution-based and Transformer-based vision backbone networks process images into the grid or sequence structures, respectively, which are inflexible for capturing irregular objects. Though Vision GNN (ViG) adopts graph-level features…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Jiafu Wu , Jian Li , Jiangning Zhang , Boshen Zhang , Mingmin Chi , Yabiao Wang , Chengjie Wang