English

LOSC: LiDAR Open-voc Segmentation Consolidator

Computer Vision and Pattern Recognition 2026-03-17 v2

Abstract

We study the use of image-based Vision-Language Models (VLMs) for open-vocabulary segmentation of lidar scans in driving settings. Classically, image semantics can be back-projected onto 3D point clouds. Yet, resulting point labels are noisy and sparse. We consolidate these labels to enforce both spatio-temporal consistency and robustness to image-level augmentations. We then train a 3D network based on these refined labels. This simple method, called LOSC, outperforms the SOTA of zero-shot open-vocabulary semantic and panoptic segmentation on both nuScenes and SemanticKITTI, with significant margins. Code is available at https://github.com/valeoai/LOSC.

Keywords

Cite

@article{arxiv.2507.07605,
  title  = {LOSC: LiDAR Open-voc Segmentation Consolidator},
  author = {Nermin Samet and Gilles Puy and Renaud Marlet},
  journal= {arXiv preprint arXiv:2507.07605},
  year   = {2026}
}

Comments

3DV 2026, Oral

R2 v1 2026-07-01T03:54:33.054Z