English
Related papers

Related papers: FMLGS: Fast Multilevel Language Embedded Gaussians…

200 papers

Learning 4D language fields to enable time-sensitive, open-ended language queries in dynamic scenes is essential for many real-world applications. While LangSplat successfully grounds CLIP features into 3D Gaussian representations,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Wanhua Li , Renping Zhou , Jiawei Zhou , Yingwei Song , Johannes Herter , Minghan Qin , Gao Huang , Hanspeter Pfister

We introduce Lifting By Gaussians (LBG), a novel approach for open-world instance segmentation of 3D Gaussian Splatted Radiance Fields (3DGS). Recently, 3DGS Fields have emerged as a highly efficient and explicit alternative to Neural…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Rohan Chacko , Nicolai Haeni , Eldar Khaliullin , Lin Sun , Douglas Lee

Reconstructing urban scenes is challenging due to their complex geometries and the presence of potentially dynamic objects. 3D Gaussian Splatting (3DGS)-based methods have shown strong performance, but existing approaches often incorporate…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Ziwen Li , Jiaxin Huang , Runnan Chen , Yunlong Che , Yandong Guo , Tongliang Liu , Fakhri Karray , Mingming Gong

3D Gaussian Splatting (3DGS), a 3D representation method with photorealistic real-time rendering capabilities, is regarded as an effective tool for narrowing the sim-to-real gap. However, it lacks fine-grained semantics and physical…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Bingchen Miao , Rong Wei , Zhiqi Ge , Xiaoquan sun , Shiqi Gao , Jingzhe Zhu , Renhan Wang , Siliang Tang , Jun Xiao , Rui Tang , Juncheng Li

This paper introduces OpenGaussian, a method based on 3D Gaussian Splatting (3DGS) capable of 3D point-level open vocabulary understanding. Our primary motivation stems from observing that existing 3DGS-based open vocabulary methods mainly…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Yanmin Wu , Jiarui Meng , Haijie Li , Chenming Wu , Yahao Shi , Xinhua Cheng , Chen Zhao , Haocheng Feng , Errui Ding , Jingdong Wang , Jian Zhang

Recent advances in leveraging large-scale Internet photo collections for 3D reconstruction have enabled immersive virtual exploration of landmarks and historic sites worldwide. However, little attention has been given to the immersive…

Graphics · Computer Science 2025-08-06 Yuze Wang , Yue Qi

Understanding open-vocabulary 3D scenes with Gaussian-based representations remains challenging due to fragmented and spatially inconsistent semantic predictions across multi-view observations. In this paper, we present OpenGaFF, a novel…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Kunyi Li , Michael Niemeyer , Sen Wang , Stefano Gasperini , Nassir Navab , Federico Tombari

Open-vocabulary 3D scene understanding enables users to segment novel objects in complex 3D environments through natural language. However, existing approaches remain slow, memory-intensive, and overly complex due to iterative optimization…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Jaehun Bang , Jinhyeok Kim , Minji Kim , Seungheon Jeong , Kyungdon Joo

Humans live in a 3D world and commonly use natural language to interact with a 3D scene. Modeling a 3D language field to support open-ended language queries in 3D has gained increasing attention recently. This paper introduces LangSplat,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Minghan Qin , Wanhua Li , Jiawei Zhou , Haoqian Wang , Hanspeter Pfister

3D affordance reasoning is essential in associating human instructions with the functional regions of 3D objects, facilitating precise, task-oriented manipulations in embodied AI. However, current methods, which predominantly depend on…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Zeming Wei , Junyi Lin , Yang Liu , Weixing Chen , Jingzhou Luo , Guanbin Li , Liang Lin

Language-guided robotic grasping is a rapidly advancing field where robots are instructed using human language to grasp specific objects. However, existing methods often depend on dense camera views and struggle to quickly update scenes,…

Robotics · Computer Science 2024-12-04 Junqiu Yu , Xinlin Ren , Yongchong Gu , Haitao Lin , Tianyu Wang , Yi Zhu , Hang Xu , Yu-Gang Jiang , Xiangyang Xue , Yanwei Fu

3D Gaussian Splatting has emerged as a powerful paradigm for explicit 3D scene representation, yet achieving efficient and consistent 3D segmentation remains challenging. Existing segmentation approaches typically rely on high-dimensional…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Wentao Sun , Quanyun Wu , Hanqing Xu , Kyle Gao , Zhengsen Xu , Yiping Chen , Dedong Zhang , Lingfei Ma , John S. Zelek , Jonathan Li

Understanding the geometric and semantic structure of environments is essential for embodied navigation and reasoning. Existing semantic mapping methods trade off between explicit geometry and multi-scale semantics, and lack a native…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Sixian Zhang , Yiyao Wang , Xinhang Song , Keming Zhang , Zijian Xu , Shuqiang Jiang

Text-to-video retrieval requires precise alignment between language and temporally rich audio-video signals. However, existing methods often emphasize visual cues while underutilizing audio semantics or relying on coarse fusion strategies,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Bowen Yang , Yun Cao , Chen He , Xiaosu Su

Holistic 3D scene understanding, which jointly models geometry, appearance, and semantics, is crucial for applications like augmented reality and robotic interaction. Existing feed-forward 3D scene understanding methods (e.g., LSM) are…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Qijing Li , Jingxiang Sun , Liang An , Zhaoqi Su , Hongwen Zhang , Yebin Liu

Building articulated objects is a key challenge in computer vision. Existing methods often fail to effectively integrate information across different object states, limiting the accuracy of part-mesh reconstruction and part dynamics…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Yu Liu , Baoxiong Jia , Ruijie Lu , Junfeng Ni , Song-Chun Zhu , Siyuan Huang

We present FLEG, a feed-forward network that reconstructs language-embedded 3D Gaussians from arbitrary views. Previous feed-forward language-embedded Gaussian reconstruction methods are restricted to a fixed number of input views and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Qijian Tian , Xin Tan , Jiayu Ying , Xuhong Wang , Yuan Xie , Lizhuang Ma

3D Gaussian Splatting (3DGS) has recently gained popularity for efficient scene rendering by representing scenes as explicit sets of anisotropic 3D Gaussians. However, most existing work focuses primarily on modeling external surfaces. In…

Image and Video Processing · Electrical Eng. & Systems 2026-01-12 Shuxin Liang , Yihan Xiao , Wenlu Tang

High-fidelity 3D reconstruction is critical for aerial inspection tasks such as infrastructure monitoring, structural assessment, and environmental surveying. While traditional photogrammetry techniques enable geometric modeling, they lack…

Graphics · Computer Science 2025-05-26 Mahmoud Chick Zaouali , Todd Charter , Homayoun Najjaran

4D millimeter-wave radar is a promising sensing modality for autonomous driving, yet effective 3D object detection from 4D radar and monocular images remains challenging. Existing fusion approaches either rely on instance proposals lacking…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Xiaokai Bai , Chenxu Zhou , Lianqing Zheng , Si-Yuan Cao , Jianan Liu , Xiaohan Zhang , Yiming Li , Zhengzhuang Zhang , Hui-liang Shen