English

SegSplat: Feed-forward Gaussian Splatting and Open-Set Semantic Segmentation

Computer Vision and Pattern Recognition 2025-11-25 v1

Abstract

We have introduced SegSplat, a novel framework designed to bridge the gap between rapid, feed-forward 3D reconstruction and rich, open-vocabulary semantic understanding. By constructing a compact semantic memory bank from multi-view 2D foundation model features and predicting discrete semantic indices alongside geometric and appearance attributes for each 3D Gaussian in a single pass, SegSplat efficiently imbues scenes with queryable semantics. Our experiments demonstrate that SegSplat achieves geometric fidelity comparable to state-of-the-art feed-forward 3D Gaussian Splatting methods while simultaneously enabling robust open-set semantic segmentation, crucially \textit{without} requiring any per-scene optimization for semantic feature integration. This work represents a significant step towards practical, on-the-fly generation of semantically aware 3D environments, vital for advancing robotic interaction, augmented reality, and other intelligent systems.

Keywords

Cite

@article{arxiv.2511.18386,
  title  = {SegSplat: Feed-forward Gaussian Splatting and Open-Set Semantic Segmentation},
  author = {Peter Siegel and Federico Tombari and Marc Pollefeys and Daniel Barath},
  journal= {arXiv preprint arXiv:2511.18386},
  year   = {2025}
}
R2 v1 2026-07-01T07:50:50.991Z