English
Related papers

Related papers: DriveVGGT: Calibration-Constrained Visual Geometry…

200 papers

Most existing 3D referring expression segmentation (3DRES) methods rely on dense, high-quality point clouds, while real-world agents such as robots and mobile phones operate with only a few sparse RGB views and strict latency constraints.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Changli Wu , Haodong Wang , Jiayi Ji , Yutian Yao , Chunsai Du , Jihua Kang , Yanwei Fu , Liujuan Cao

Driven by the advancement of 3D devices, stereo vision tasks including stereo matching and stereo conversion have emerged as a critical research frontier. Contemporary stereo vision backbones typically rely on either monocular depth…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Ziyang Chen , Yansong Qu , You Shen , Xuan Cheng , Liujuan Cao

This paper presents VGGT-360, a novel training-free framework for zero-shot, geometry-consistent panoramic depth estimation. Unlike prior view-independent training-free approaches, VGGT-360 reformulates the task as panoramic reprojection…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Jiayi Yuan , Haobo Jiang , De Wen Soh , Na Zhao

3D instance segmentation methods typically rely on high-quality point clouds or posed RGB-D scans, requiring complex multi-stage processing pipelines, and are highly sensitive to reconstruction noise. While recent feed-forward transformers…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Jinyuan Qu , Hongyang Li , Lei Zhang

Visual Geometry Transformer (VGGT) is a strong feed-forward model for multiple 3D tasks, but its Alternating-Attention (AA) stack scales quadratically in the total token count, making long clips expensive. Existing token-reduction…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Haotang Li , Zhenyu Qi , Shaohan Henry Wang , Kebin Peng , Zi Wang , Qing Guo , Sen He , Huanrui Yang

Recent feed-forward reconstruction models, such as VGGT, have proven competitive with traditional optimization-based reconstructors while also providing geometry-aware features useful for other tasks. Here, we show that the quality of these…

Dense visual odometry (VO), which provides pose estimation and dense 3D reconstruction, serves as the cornerstone for applications ranging from robotics to augmented reality. Recently, feed-forward models have demonstrated remarkable…

Robotics · Computer Science 2026-04-03 Junxiang Pan , Lipu Zhou , Baojie Chen

Visual Geometry Grounded Transformer (VGGT) advances 3D reconstruction via scalable Transformer architecture, but the quadratic complexity of global attention prevents long context application. StreamVGGT enables streaming with causal…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Zichen Zou , Xiaosong Jia , Zuxuan Wu , Yu-Gang Jiang

Reconstructing dynamic 4D scenes is challenging, as it requires robust disentanglement of dynamic objects from the static background. While 3D foundation models like VGGT provide accurate 3D geometry, their performance drops markedly when…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Yu Hu , Chong Cheng , Sicheng Yu , Xiaoyang Guo , Hao Wang

Recent 3D feed-forward models, such as the Visual Geometry Grounded Transformer (VGGT), have shown strong capability in inferring 3D attributes of static scenes. However, since they are typically trained on static datasets, these models…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Kaichen Zhou , Yuhan Wang , Grace Chen , Xinhai Chang , Gaspard Beaudouin , Fangneng Zhan , Paul Pu Liang , Mengyu Wang

Models such as VGGT and $\pi^3$ have shown strong multi-view 3D performance, but their heavy reliance on global self-attention results in high computational cost. Existing sparse-attention variants offer partial speedups, yet lack a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Xianbing Sun , Zhikai Zhu , Zhengyu Lou , Bo Yang , Jinyang Tang , Liqing Zhang , He Wang , Jianfu Zhang

3D reconstruction in large-scale scenes is a fundamental task in 3D perception, but the inherent trade-off between accuracy and computational efficiency remains a significant challenge. Existing methods either prioritize speed and produce…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Jungho Lee , Minhyeok Lee , Sunghun Yang , Minseok Kang , Sangyoun Lee

Although existing monocular depth estimation methods have made great progress, predicting an accurate absolute depth map from a single image is still challenging due to the limited modeling capacity of networks and the scale ambiguity…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Jie Xiang , Yun Wang , Lifeng An , Haiyang Liu , Zijun Wang , Jian Liu

Reconstructing dynamic 4D scenes from monocular videos is a fundamental yet challenging task. While recent 3D foundation models provide strong geometric priors, their performance significantly degrades in dynamic environments. This…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Ying Zang , Xuanyi Liu , Yidong Han , Deyi Ji , Chaotao Ding , Yuanqi Hu , Qi Zhu , Xuanfu Li , Jin Ma , Lingyun Sun , Tianrun Chen , Lanyun Zhu

The grand vision of enabling persistent, large-scale 3D visual geometry understanding is shackled by the irreconcilable demands of scalability and long-term stability. While offline models like VGGT achieve inspiring geometry capability,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Shuai Yuan , Yantai Yang , Xiaotian Yang , Xupeng Zhang , Zhonghao Zhao , Lingming Zhang , Zhipeng Zhang

Visual imitation learning frameworks allow robots to learn manipulation skills from expert demonstrations. While existing approaches mainly focus on policy design, they often neglect the structure and capacity of visual encoders, limiting…

Robotics · Computer Science 2025-09-24 Shijia Ge , Yinxin Zhang , Shuzhao Xie , Weixiang Zhang , Mingcai Zhou , Zhi Wang

Despite recent progress in calibration-free monocular SLAM via 3D vision foundation models, scale drift remains severe on long sequences. Motion-agnostic partitioning breaks contextual coherence and causes zero-motion drift, while…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Zhuang Xiong , Chen Zhang , Qingshan Xu , Wenbing Tao

State-of-the-art 3D computer vision algorithms continue to advance in handling sparse, unordered image sets. Recently developed foundational models for 3D reconstruction, such as Dense and Unconstrained Stereo 3D Reconstruction (DUSt3R),…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Xinyi Wu , Steven Landgraf , Markus Ulrich , Rongjun Qin

Volumetric visualization has long been dominated by Direct Volume Rendering (DVR), which operates on dense voxel grids and suffers from limited scalability as resolution and interactivity demands increase. Recent advances in 3D Gaussian…

Graphics · Computer Science 2026-04-15 Yuxuan Wang , Qibiao Li , Youcheng Cai

Reconstructing dynamic 4D scenes is an important yet challenging task. While 3D foundation models like VGGT excel in static settings, they often struggle with dynamic sequences where motion causes significant geometric ambiguity. To address…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Ying Zang , Yidong Han , Chaotao Ding , Yuanqi Hu , Deyi Ji , Qi Zhu , Xuanfu Li , Jin Ma , Lingyun Sun , Tianrun Chen , Lanyun Zhu