English
Related papers

Related papers: 4D LangSplat: 4D Language Gaussian Splatting via M…

200 papers

Egocentric scenes exhibit frequent occlusions, varied viewpoints, and dynamic interactions compared to typical scene understanding tasks. Occlusions and varied viewpoints can lead to multi-view semantic inconsistencies, while dynamic…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Di Li , Jie Feng , Jiahao Chen , Weisheng Dong , Guanbin Li , Guangming Shi , Licheng Jiao

Understanding dynamic 4D environments through natural language queries requires not only accurate scene reconstruction but also robust semantic grounding across space, time, and viewpoints. While recent methods using neural representations…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Ruilin Tang , Yang Zhou , Zhong Ye , Wenxi Liu , Yan Huang , Shengfeng He

To enable AI agents to interact seamlessly with both humans and 3D environments, they must not only perceive the 3D world accurately but also align human language with 3D spatial representations. While prior work has made significant…

Artificial Intelligence · Computer Science 2025-09-26 Saimouli Katragadda , Cho-Ying Wu , Yuliang Guo , Xinyu Huang , Guoquan Huang , Liu Ren

Realistic reconstruction of dynamic 4D scenes from monocular videos is essential for understanding the physical world. Despite recent progress in neural rendering, existing methods often struggle to recover accurate 3D geometry and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Haoran Zhou , Gim Hee Lee

Recent advances in driving-scene generation and reconstruction have demonstrated significant potential for enhancing autonomous driving systems by producing scalable and controllable training data. Existing generation methods primarily…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Ziyue Zhu , Zhanqian Wu , Zhenxin Zhu , Lijun Zhou , Haiyang Sun , Bing Wan , Kun Ma , Guang Chen , Hangjun Ye , Jin Xie , jian Yang

Capturing 4D spatiotemporal surroundings is crucial for the safe and reliable operation of robots in dynamic environments. However, most existing methods address only one side of the problem: they either provide coarse geometric tracking…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Maximilian Luz , Rohit Mohan , Thomas Nürnberg , Yakov Miron , Daniele Cattaneo , Abhinav Valada

Language-guided 3D scene understanding is important for advancing applications in robotics, AR/VR, and human-computer interaction, enabling models to comprehend and interact with 3D environments through natural language. While 2D…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Anh Thai , Songyou Peng , Kyle Genova , Leonidas Guibas , Thomas Funkhouser

Previous text-to-4D methods have leveraged multiple Score Distillation Sampling (SDS) techniques, combining motion priors from video-based diffusion models (DMs) with geometric priors from multiview DMs to implicitly guide 4D renderings.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-20 Qiaowei Miao , JinSheng Quan , Kehan Li , Yawei Luo

Constructing 4D language fields is crucial for embodied AI, augmented/virtual reality, and 4D scene understanding, as they provide enriched semantic representations of dynamic environments and enable open-vocabulary querying in complex…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Xianfeng Wu , Yajing Bai , Minghan Li , Xianzu Wu , Xueqi Zhao , Zhongyuan Lai , Wenyu Liu , Xinggang Wang

LongSplat addresses critical challenges in novel view synthesis (NVS) from casually captured long videos characterized by irregular camera motion, unknown camera poses, and expansive scenes. Current methods often suffer from pose drift,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Chin-Yang Lin , Cheng Sun , Fu-En Yang , Min-Hung Chen , Yen-Yu Lin , Yu-Lun Liu

Representing and rendering dynamic scenes has been an important but challenging task. Especially, to accurately model complex motions, high efficiency is usually hard to guarantee. To achieve real-time dynamic scene rendering while also…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Guanjun Wu , Taoran Yi , Jiemin Fang , Lingxi Xie , Xiaopeng Zhang , Wei Wei , Wenyu Liu , Qi Tian , Xinggang Wang

In robot-assisted minimally invasive surgery, high-fidelity dynamic endoscopic scene reconstruction and simulation are crucial to enhancing downstream tasks and advancing surgical outcomes. However, existing methods primarily focus on…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Changjing Liu , Yiming Huang , Long Bai , Beilei Cui , Hongliang Ren

Modeling dynamic scenes through 4D Gaussians offers high visual fidelity and fast rendering speeds, but comes with significant storage overhead. Recent approaches mitigate this cost by aggressively reducing the number of Gaussians. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Woong Oh Cho , In Cho , Seoha Kim , Jeongmin Bae , Youngjung Uh , Seon Joo Kim

Open-vocabulary 3D scene understanding presents a significant challenge in computer vision, with wide-ranging applications in embodied agents and augmented reality systems. Existing methods adopt neurel rendering methods as 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Jun Guo , Xiaojian Ma , Yue Fan , Huaping Liu , Qing Li

We introduce Dr. Splat, a novel approach for open-vocabulary 3D scene understanding leveraging 3D Gaussian Splatting. Unlike existing language-embedded 3DGS methods, which rely on a rendering process, our method directly associates…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Kim Jun-Seong , GeonU Kim , Kim Yu-Ji , Yu-Chiang Frank Wang , Jaesung Choe , Tae-Hyun Oh

Language-augmented scene representations hold great promise for large-scale robotics applications such as search-and-rescue, smart cities, and mining. Many of these scenarios are time-sensitive, requiring rapid scene encoding while also…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Laszlo Szilagyi , Francis Engelmann , Jeannette Bohg

3D Gaussian Splatting (3DGS) has emerged as a powerful representation for neural scene reconstruction, offering high-quality novel view synthesis while maintaining computational efficiency. In this paper, we extend the capabilities of 3DGS…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Jens Piekenbrinck , Christian Schmidt , Alexander Hermans , Narunas Vaskevicius , Timm Linder , Bastian Leibe

Novel view synthesis has shown rapid progress recently, with methods capable of producing increasingly photorealistic results. 3D Gaussian Splatting has emerged as a promising method, producing high-quality renderings of scenes and enabling…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Richard Shaw , Michal Nazarczuk , Jifei Song , Arthur Moreau , Sibi Catley-Chandar , Helisa Dhamo , Eduardo Perez-Pellitero

3D Gaussian Splatting (3DGS) has garnered significant attention due to its superior scene representation fidelity and real-time rendering performance, especially for dynamic 3D scene reconstruction (\textit{i.e.}, 4D reconstruction).…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Henan Wang , Hanxin Zhu , Xinliang Gong , Tianyu He , Xin Li , Zhibo Chen

Language-driven 3D Gaussian Splatting (3DGS) editing provides a more convenient approach for modifying complex scenes in VR/AR. Standard pipelines typically adopt a two-stage strategy: first editing multiple 2D views, and then optimizing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Yanhui Chen , Jiahong Li , Jingchao Wang , Junyi Lin , Zixin Zeng , Yang Shi