English

Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction

Computer Vision and Pattern Recognition 2025-08-20 v2

Abstract

We introduce Geo4D, a method to repurpose video diffusion models for monocular 3D reconstruction of dynamic scenes. By leveraging the strong dynamic priors captured by large-scale pre-trained video models, Geo4D can be trained using only synthetic data while generalizing well to real data in a zero-shot manner. Geo4D predicts several complementary geometric modalities, namely point, disparity, and ray maps. We propose a new multi-modal alignment algorithm to align and fuse these modalities, as well as a sliding window approach at inference time, thus enabling robust and accurate 4D reconstruction of long videos. Extensive experiments across multiple benchmarks show that Geo4D significantly surpasses state-of-the-art video depth estimation methods.

Keywords

Cite

@article{arxiv.2504.07961,
  title  = {Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction},
  author = {Zeren Jiang and Chuanxia Zheng and Iro Laina and Diane Larlus and Andrea Vedaldi},
  journal= {arXiv preprint arXiv:2504.07961},
  year   = {2025}
}

Comments

17 pages, 6 figures, ICCV 2025 Highlight, Project page: https://geo4d.github.io/