English

VIDAR: Visual-Inertial Dense Alignment and Reconstruction via a Geometric Foundation Model

Robotics 2026-07-19 v1

Abstract

Monocular foundation models provide dense geometry but usually lack a stable metric scale. This paper presents VIDAR, a visual-inertial dense reconstruction framework that couples SVO+IMU odometry with Depth Anything 3. VIDAR uses the visual-inertial front end as a metric anchor: it provides camera poses, scale, and a consistent world frame for aligning dense foundation-model predictions across time. The foundation model then contributes detailed local geometry that is fused into a global reconstruction. We study both pose-conditioned DA3 and a decoupled alignment strategy. On EuRoC, pose injection reduces scale error to about 1\% and reaches 0.463 mean [email protected]; the decoupled hybrid improves this to 0.676 without ground-truth poses. Results on EuRoC and TUM RGB-D show that VIDAR is a practical route to metric dense monocular reconstruction.

Keywords

Cite

@article{arxiv.2607.17171,
  title  = {VIDAR: Visual-Inertial Dense Alignment and Reconstruction via a Geometric Foundation Model},
  author = {Diyari Mohammed Salih and Lingxiang Hu and Naima AitOufroukh-Mammar and Fabien Bonardi},
  journal= {arXiv preprint arXiv:2607.17171},
  year   = {2026}
}