English

Geospatial-Prior Guidance for 3D Semantic Scene Completion

Computer Vision and Pattern Recognition 2026-08-04 v1

Abstract

Inferring complete 3D geometry and semantics from onboard images remains challenging because occlusions and restricted fields of view leave large scene regions underconstrained. Although satellite imagery provides wide-area context, appearance cues alone offer limited structural guidance and may be unreliable because of spatial or temporal discrepancies. We present GeoScene, a geospatially guided framework that jointly uses satellite imagery and structured OpenStreetMap cues as soft priors for 3D semantic scene completion. GeoScene learns complementary voxel-wise reliability weights for onboard observations and geospatial guidance, and uses them to control feature refinement in observed and unobserved regions. This design preserves local visual evidence while exploiting large-scale road and building structure beyond onboard visibility. Experiments on SemanticKITTI and SSCBench-KITTI-360 demonstrate that GeoScene consistently improves both geometric and semantic completion under the geospatial-prior-assisted setting, with the most pronounced benefits for large-scale static and geospatially structured classes.

Cite

@article{arxiv.2608.03618,
  title  = {Geospatial-Prior Guidance for 3D Semantic Scene Completion},
  author = {Meng Wang and Shougao Zhang and Wenzhe He and Ruihui Li and Nan Hu and Zhuo Tang and Kenli Li},
  journal= {arXiv preprint arXiv:2608.03618},
  year   = {2026}
}