English

LangFlash: Feed-forward 3D Language Gaussian Splatting from Sparse Unposed Images

Computer Vision and Pattern Recognition 2026-05-25 v1

Abstract

We present LangFlash, a feed-forward framework for 3D Language Gaussian Splatting that reconstructs 3D scenes parameterized by Gaussian primitives enriched with language-aligned semantic features from sparse unposed multi-view images. Unlike optimization-based 3D methods, LangFlash directly predicts the geometry and semantics in a single forward pass, enabling low-latency 3D reconstruction and language-consistent scene understanding. To support large-scale training, we enriched the RealEstate10k dataset with coherent and dense semantic information for 3D semantic supervision. Furthermore, we propose a sparse semantic encoding scheme that combines a global semantic dictionary with locally varying per-primitive weights, preserving high-level linguistic information, while reducing representation complexity. Experimental results show that LangFlash achieves superior novel view synthesis and semantic consistency compared with previous methods. This study establishes a new paradigm for pose-free, language-grounded 3D scene reconstruction, advancing generalizable 3D vision and multimodal scene understanding. Demo is available at https://liylo.github.io/langflash.github.io/.

Keywords

Cite

@article{arxiv.2605.23287,
  title  = {LangFlash: Feed-forward 3D Language Gaussian Splatting from Sparse Unposed Images},
  author = {Yilong Liu and Wanhua Li and Chen Zhu-Tian and Hanspeter Pfister},
  journal= {arXiv preprint arXiv:2605.23287},
  year   = {2026}
}

Comments

CVPRF 2026

R2 v1 2026-07-22T07:27:42.319Z