English

From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs

Computer Vision and Pattern Recognition 2025-06-10 v2

Abstract

3D vision-language grounding faces a fundamental data bottleneck: while 2D models train on billions of images, 3D models have access to only thousands of labeled scenes--a six-order-of-magnitude gap that severely limits performance. We introduce LIFT-GS\textbf{LIFT-GS}, a practical distillation technique that overcomes this limitation by using differentiable rendering to bridge 3D and 2D supervision. LIFT-GS predicts 3D Gaussian representations from point clouds and uses them to render predicted language-conditioned 3D masks into 2D views, enabling supervision from 2D foundation models (SAM, CLIP, LLaMA) without requiring any 3D annotations. This render-supervised formulation enables end-to-end training of complete encoder-decoder architectures and is inherently model-agnostic. LIFT-GS achieves state-of-the-art results with 25.7%25.7\% mAP on open-vocabulary instance segmentation (vs. 20.2%20.2\% prior SOTA) and consistent 1030%10-30\% improvements on referential grounding tasks. Remarkably, pretraining effectively multiplies fine-tuning datasets by 2X, demonstrating strong scaling properties that suggest 3D VLG currently operates in a severely data-scarce regime. Project page: https://liftgs.github.io

Keywords

Cite

@article{arxiv.2502.20389,
  title  = {From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs},
  author = {Ang Cao and Sergio Arnaud and Oleksandr Maksymets and Jianing Yang and Ayush Jain and Sriram Yenamandra and Ada Martin and Vincent-Pierre Berges and Paul McVay and Ruslan Partsey and Aravind Rajeswaran and Franziska Meier and Justin Johnson and Jeong Joon Park and Alexander Sax},
  journal= {arXiv preprint arXiv:2502.20389},
  year   = {2025}
}

Comments

Project page: https://liftgs.github.io

R2 v1 2026-06-28T22:00:39.839Z