English
Related papers

Related papers: DVGT: Driving Visual Geometry Transformer

200 papers

An accurate understanding of a self-driving vehicle's surrounding environment is crucial for its navigation system. To enhance the effectiveness of existing algorithms and facilitate further research, it is essential to provide…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Abtin Mahyar , Hossein Motamednia , Dara Rahmati

Dense visual odometry (VO), which provides pose estimation and dense 3D reconstruction, serves as the cornerstone for applications ranging from robotics to augmented reality. Recently, feed-forward models have demonstrated remarkable…

Robotics · Computer Science 2026-04-03 Junxiang Pan , Lipu Zhou , Baojie Chen

Current methods for dense 3D point tracking in dynamic scenes typically rely on pairwise processing, require known camera poses, or assume temporal ordering of input frames, thereby constraining their flexibility and applicability.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Vivek Alumootil , Tuan-Anh Vu

Understanding and reconstructing the complex geometry and motion of dynamic scenes from video remains a formidable challenge in computer vision. This paper introduces D4RT, a simple yet powerful feedforward model designed to efficiently…

World model-based searching and planning are widely recognized as a promising path toward human-level physical intelligence. However, current driving world models primarily rely on video diffusion models, which specialize in visual…

Computer Vision and Pattern Recognition · Computer Science 2024-12-25 Yuntao Chen , Yuqi Wang , Zhaoxiang Zhang

Cross-view video geo-localization (CVGL) aims to derive GPS trajectories from street-view videos by aligning them with aerial-view images. Despite their promising performance, current CVGL methods face significant challenges. These methods…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Manu S Pillai , Mamshad Nayeem Rizve , Mubarak Shah

Perceiving and reconstructing 3D geometry from videos is a fundamental yet challenging computer vision task. To facilitate interactive and low-latency applications, we propose a streaming visual geometry transformer that shares a similar…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Dong Zhuo , Wenzhao Zheng , Jiahe Guo , Yuqi Wu , Jie Zhou , Jiwen Lu

We study a crucial yet often overlooked issue inherent to Vision Transformers (ViTs): feature maps of these models exhibit grid-like artifacts, which hurt the performance of ViTs in downstream dense prediction tasks such as semantic…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Jiawei Yang , Katie Z Luo , Jiefeng Li , Congyue Deng , Leonidas Guibas , Dilip Krishnan , Kilian Q Weinberger , Yonglong Tian , Yue Wang

Reconstructing dynamic 4D scenes is an important yet challenging task. While 3D foundation models like VGGT excel in static settings, they often struggle with dynamic sequences where motion causes significant geometric ambiguity. To address…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Ying Zang , Yidong Han , Chaotao Ding , Yuanqi Hu , Deyi Ji , Qi Zhu , Xuanfu Li , Jin Ma , Lingyun Sun , Tianrun Chen , Lanyun Zhu

This paper presents VGGT-360, a novel training-free framework for zero-shot, geometry-consistent panoramic depth estimation. Unlike prior view-independent training-free approaches, VGGT-360 reformulates the task as panoramic reprojection…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Jiayi Yuan , Haobo Jiang , De Wen Soh , Na Zhao

Dynamic Gaussian splatting has led to impressive scene reconstruction and image synthesis advances in novel views. Existing methods, however, heavily rely on pre-computed poses and Gaussian initialization by Structure from Motion (SfM)…

Computer Vision and Pattern Recognition · Computer Science 2024-06-27 Hao Li , Jingfeng Li , Dingwen Zhang , Chenming Wu , Jieqi Shi , Chen Zhao , Haocheng Feng , Errui Ding , Jingdong Wang , Junwei Han

This paper proposes 3DGeoDet, a novel geometry-aware 3D object detection approach that effectively handles single- and multi-view RGB images in indoor and outdoor environments, showcasing its general-purpose applicability. The key challenge…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Yi Zhang , Yi Wang , Yawen Cui , Lap-Pui Chau

Recent advancements in world models have revolutionized dynamic environment simulation, allowing systems to foresee future states and assess potential actions. In autonomous driving, these capabilities help vehicles anticipate the behavior…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Anthony Chen , Wenzhao Zheng , Yida Wang , Xueyang Zhang , Kun Zhan , Peng Jia , Kurt Keutzer , Shanghang Zhang

Realistic speech-driven 3D facial animation is a challenging problem due to the complex relationship between speech and face. In this paper, we propose a deep architecture, called Geometry-guided Dense Perspective Network (GDPnet), to…

Graphics · Computer Science 2020-08-25 Jingying Liu , Binyuan Hui , Kun Li , Yunke Liu , Yu-Kun Lai , Yuxiang Zhang , Yebin Liu , Jingyu Yang

Holistically understanding an object and its 3D movable parts through visual perception models is essential for enabling an autonomous agent to interact with the world. For autonomous driving, the dynamics and states of vehicle parts such…

Computer Vision and Pattern Recognition · Computer Science 2021-01-07 Feixiang Lu , Zongdai Liu , Hui Miao , Peng Wang , Liangjun Zhang , Ruigang Yang , Dinesh Manocha , Bin Zhou

Generating accurate 3D models is a challenging problem that traditionally requires explicit learning from 3D datasets using supervised learning. Although recent advances have shown promise in learning 3D models from 2D images, these methods…

Computer Vision and Pattern Recognition · Computer Science 2024-02-05 Qijia Shen , Guangrun Wang

Integrating open-vocabulary semantic information into dynamic 3D scene representations is essential for long-term embodied scene understanding. However, existing methods often suffer from fragile instance association due to incomplete…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Luzhou Ge , Xiangyu Zhu , Jinyan Liu , Xuesong Li

We tackle the task of synthesizing novel views of an object given a few input images and associated camera viewpoints. Our work is inspired by recent 'geometry-free' approaches where multi-view images are encoded as a (global) set-latent…

Computer Vision and Pattern Recognition · Computer Science 2023-01-12 Naveen Venkat , Mayank Agarwal , Maneesh Singh , Shubham Tulsiani

Synthesis of diverse driving scenes serves as a crucial data augmentation technique for validating the robustness and generalizability of autonomous driving systems. Current methods aggregate high-definition (HD) maps and 3D bounding boxes…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Zhechao Wang , Yiming Zeng , Lufan Ma , Zeqing Fu , Chen Bai , Ziyao Lin , Cheng Lu

Predominant techniques on talking head generation largely depend on 2D information, including facial appearances and motions from input face images. Nevertheless, dense 3D facial geometry, such as pixel-wise depth, plays a critical role in…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Fa-Ting Hong , Li Shen , Dan Xu