English

NVSMask3D: Hard Visual Prompting with Camera Pose Interpolation for 3D Open Vocabulary Instance Segmentation

Computer Vision and Pattern Recognition 2025-04-22 v1

Abstract

Vision-language models (VLMs) have demonstrated impressive zero-shot transfer capabilities in image-level visual perception tasks. However, they fall short in 3D instance-level segmentation tasks that require accurate localization and recognition of individual objects. To bridge this gap, we introduce a novel 3D Gaussian Splatting based hard visual prompting approach that leverages camera interpolation to generate diverse viewpoints around target objects without any 2D-3D optimization or fine-tuning. Our method simulates realistic 3D perspectives, effectively augmenting existing hard visual prompts by enforcing geometric consistency across viewpoints. This training-free strategy seamlessly integrates with prior hard visual prompts, enriching object-descriptive features and enabling VLMs to achieve more robust and accurate 3D instance segmentation in diverse 3D scenes.

Keywords

Cite

@article{arxiv.2504.14638,
  title  = {NVSMask3D: Hard Visual Prompting with Camera Pose Interpolation for 3D Open Vocabulary Instance Segmentation},
  author = {Junyuan Fang and Zihan Wang and Yejun Zhang and Shuzhe Wang and Iaroslav Melekhov and Juho Kannala},
  journal= {arXiv preprint arXiv:2504.14638},
  year   = {2025}
}

Comments

15 pages, 4 figures, Scandinavian Conference on Image Analysis 2025

R2 v1 2026-06-28T23:04:47.357Z