English

Joint 2D-3D Segmentation and Association in Street-level Imaging

Computer Vision and Pattern Recognition 2026-05-27 v1

Abstract

Accurate interpretation of street-level imagery is essential for large-scale urban mapping and the creation of Spatial Digital Twin (SDT) environments. This work presents a unified framework for joint 2D-3D segmentation and association that integrates visual semantics with multi-view geometric reasoning. Unlike conventional approaches that rely heavily on sequential frames for temporal tracking, our method leverages zero-shot detection and segmentation together with structure-from-motion reconstruction to establish stable cross-view correspondences. A 3D-driven association mechanism replaces traditional 2D multi-object tracking, using geometric consistency to guide identity preservation across wide-baseline viewpoints and varying imaging conditions. By combining 2D texture cues with global 3D context, the proposed pipeline is well-suited for scalable street-level processing and can be used for a variety of object types. Experiments demonstrate substantially improved coverage of ground-truth sequences and more robust identity retention compared to state-of-the-art 2D-only tracking methods, achieving a 22% performance gain in challenging urban scenarios.

Keywords

Cite

@article{arxiv.2605.26725,
  title  = {Joint 2D-3D Segmentation and Association in Street-level Imaging},
  author = {Amir Melnikov and Masayuki Tanaka and Yusuke Monno and Masatoshi Okutomi},
  journal= {arXiv preprint arXiv:2605.26725},
  year   = {2026}
}

Comments

15 pages, 6 image figures, 1 in-body table, 1 in-body algorithm, 2 indexes with tables

R2 v1 2026-07-22T07:34:08.107Z