English

AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian Splatting

Computer Vision and Pattern Recognition 2025-08-05 v3

Abstract

Obtaining high-quality 3D semantic occupancy from raw sensor data remains an essential yet challenging task, often requiring extensive manual labeling. In this work, we propose AutoOcc, a vision-centric automated pipeline for open-ended semantic occupancy annotation that integrates differentiable Gaussian splatting guided by vision-language models. We formulate the open-ended semantic 3D occupancy reconstruction task to automatically generate scene occupancy by combining attention maps from vision-language models and foundation vision models. We devise semantic-aware Gaussians as intermediate geometric descriptors and propose a cumulative Gaussian-to-voxel splatting algorithm that enables effective and efficient occupancy annotation. Our framework outperforms existing automated occupancy annotation methods without human labels. AutoOcc also enables open-ended semantic occupancy auto-labeling, achieving robust performance in both static and dynamically complex scenarios.

Keywords

Cite

@article{arxiv.2502.04981,
  title  = {AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian Splatting},
  author = {Xiaoyu Zhou and Jingqi Wang and Yongtao Wang and Yufei Wei and Nan Dong and Ming-Hsuan Yang},
  journal= {arXiv preprint arXiv:2502.04981},
  year   = {2025}
}

Comments

ICCV 2025 Hightlight (main conference)

R2 v1 2026-06-28T21:36:11.908Z