English
Related papers

Related papers: 4D LangSplat: 4D Language Gaussian Splatting via M…

200 papers

Humans excel at forecasting the future dynamics of a scene given just a single image. Video generation models that can mimic this ability are an essential component for intelligent systems. Recent approaches have improved temporal coherence…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Melonie de Almeida , Daniela Ivanova , Tong Shi , John H. Williamson , Paul Henderson

Dynamic 4D Gaussian Splatting (4DGS) effectively extends the high-speed rendering capabilities of 3D Gaussian Splatting (3DGS) to represent volumetric videos. However, the large number of Gaussians, substantial temporal redundancies, and…

Graphics · Computer Science 2026-01-14 Hyeongmin Lee , Kyungjune Baek

The efficient spatial allocation of primitives serves as the foundation of 3D Gaussian Splatting, as it directly dictates the synergy between representation compactness, reconstruction speed, and rendering fidelity. Previous solutions,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Roni Itkin , Noam Issachar , Yehonatan Keypur , Xingyu Chen , Anpei Chen , Sagie Benaim

Open-vocabulary querying in 3D space is challenging but essential for scene understanding tasks such as object localization and segmentation. Language-embedded scene representations have made progress by incorporating language features into…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Jin-Chuan Shi , Miao Wang , Hao-Bin Duan , Shao-Hua Guan

Due to the complex and highly dynamic motions in the real world, synthesizing dynamic videos from multi-view inputs for arbitrary viewpoints is challenging. Previous works based on neural radiance field or 3D Gaussian splatting are limited…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Jiahao Wu , Rui Peng , Jianbo Jiao , Jiayu Yang , Luyang Tang , Kaiqiang Xiong , Jie Liang , Jinbo Yan , Runling Liu , Ronggang Wang

We present latentSplat, a method to predict semantic Gaussians in a 3D latent space that can be splatted and decoded by a light-weight generative 2D architecture. Existing methods for generalizable 3D reconstruction either do not scale to…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Christopher Wewer , Kevin Raj , Eddy Ilg , Bernt Schiele , Jan Eric Lenssen

Scene representations using 3D Gaussian primitives have produced excellent results in modeling the appearance of static and dynamic 3D scenes. Many graphics applications, however, demand the ability to manipulate both the appearance and the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Ri-Zhao Qiu , Ge Yang , Weijia Zeng , Xiaolong Wang

Multimodal Large Language Models (MLLMs) exhibit impressive capabilities across a variety of tasks, especially when equipped with carefully designed visual prompts. However, existing studies primarily focus on logical reasoning and visual…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Dingning Liu , Cheng Wang , Peng Gao , Renrui Zhang , Xinzhu Ma , Yuan Meng , Zhihui Wang

Vision language models (VLMs) have shown remarkable capabilities in integrating linguistic and visual reasoning but remain fundamentally limited in understanding dynamic spatiotemporal interactions. Humans effortlessly track and reason…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Shijie Zhou , Alexander Vilesov , Xuehai He , Ziyu Wan , Shuwang Zhang , Aditya Nagachandra , Di Chang , Dongdong Chen , Xin Eric Wang , Achuta Kadambi

Lifting 2D open-vocabulary understanding into 3D Gaussian Splatting (3DGS) scenes is a critical challenge. Mainstream methods, built on an embedding paradigm, suffer from three key flaws: (i) geometry-semantic inconsistency, where points,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Jiayu Ding , Xinpeng Liu , Zhiyi Pan , Shiqiang Long , Ge Li

We propose VideoRFSplat, a direct text-to-3D model leveraging a video generation model to generate realistic 3D Gaussian Splatting (3DGS) for unbounded real-world scenes. To generate diverse camera poses and unbounded spatial extent of…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Hyojun Go , Byeongjun Park , Hyelin Nam , Byung-Hoon Kim , Hyungjin Chung , Changick Kim

Open-vocabulary 3D scene understanding is crucial for applications requiring natural language-driven spatial interpretation, such as robotics and augmented reality. While 3D Gaussian Splatting (3DGS) offers a powerful representation for…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Wei Sun , Yanzhao Zhou , Jianbin Jiao , Yuan Li

Handling the dynamic environments is a significant research challenge in Visual Simultaneous Localization and Mapping (SLAM). Recent research combines 3D Gaussian Splatting (3DGS) with SLAM to achieve both robust camera pose estimation and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Yunsong Wang , Gim Hee Lee

Reconstructing large-scale dynamic driving scenes remains challenging due to the coexistence of static environments with extreme depth variation and diverse dynamic actors exhibiting complex motions. Existing Gaussian Splatting based…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Cong Wang , Ruiqi Song , Wei Tian , Chenming Zhang , Lingxi Li , Long Chen

Recent advancements in 3D Gaussian Splatting (3DGS) have demonstrated its potential for efficient and photorealistic 3D reconstructions, which is crucial for diverse applications such as robotics and immersive media. However, current…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Daheng Yin , Isaac Ding , Yili Jin , Jianxin Shi , Jiangchuan Liu

Recently, Gaussian Splatting methods have emerged as a desirable substitute for prior Radiance Field methods for novel-view synthesis of scenes captured with multi-view images or videos. In this work, we propose a novel extension to 4D…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Karly Hou , Wanhua Li , Hanspeter Pfister

Recent advancements in foundation models for 2D vision have substantially improved the analysis of dynamic scenes from monocular videos. However, despite their strong generalization capabilities, these models often lack 3D consistency, a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Haoran Zhou , Gim Hee Lee

In real-world scenarios, environment changes caused by human or agent activities make it extremely challenging for robots to perform various long-term tasks. Recent works typically struggle to effectively understand and adapt to dynamic…

Robotics · Computer Science 2025-12-19 Luzhou Ge , Xiangyu Zhu , Zhuo Yang , Xuesong Li

3D Gaussian Splatting (3DGS) has emerged as a novel explicit representation for 3D scenes, offering both high-fidelity reconstruction and efficient rendering. However, 3DGS lacks 3D segmentation ability, which limits its applicability in…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Yupeng Zhang , Dezhi Zheng , Ping Lu , Han Zhang , Lei Wang , Liping xiang , Cheng Luo , Kaijun Deng , Xiaowen Fu , Linlin Shen , Jinbao Wang

Recent advancements in multi-modal large language models (MLLMs) have shown strong potential for 3D scene understanding. However, existing methods struggle with fine-grained object grounding and contextual reasoning, limiting their ability…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Haifeng Huang , Yilun Chen , Zehan Wang , Jiangmiao Pang , Zhou Zhao
‹ Prev 1 3 4 5 6 7 10 Next ›