中文
相关论文

相关论文: SplatTalk: 3D VQA with Gaussian Splatting

200 篇论文

3D Gaussian Splatting is renowned for its high-fidelity reconstructions and real-time novel view synthesis, yet its lack of semantic understanding limits object-level perception. In this work, we propose ObjectGS, an object-aware framework…

图形学 · 计算机科学 2025-07-22 Ruijie Zhu , Mulin Yu , Linning Xu , Lihan Jiang , Yixuan Li , Tianzhu Zhang , Jiangmiao Pang , Bo Dai

Recent advancements in dynamic 3D scene reconstruction have shown promising results, enabling high-fidelity 3D novel view synthesis with improved temporal consistency. Among these, 4D Gaussian Splatting (4DGS) has emerged as an appealing…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Seungjun Oh , Younggeun Lee , Hyejin Jeon , Eunbyung Park

We introduce ShelfGaussian, an open-vocabulary multi-modal Gaussian-based 3D scene understanding framework supervised by off-the-shelf vision foundation models (VFMs). Gaussian-based methods have demonstrated superior performance and…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Lingjun Zhao , Yandong Luo , James Hays , Lu Gan

Reliable multimodal sensor fusion algorithms require accurate spatiotemporal calibration. Recently, targetless calibration techniques based on implicit neural representations have proven to provide precise and robust results. Nevertheless,…

计算机视觉与模式识别 · 计算机科学 2024-09-19 Quentin Herau , Moussab Bennehar , Arthur Moreau , Nathan Piasco , Luis Roldao , Dzmitry Tsishkou , Cyrille Migniot , Pascal Vasseur , Cédric Demonceaux

Simultaneous Localization and Mapping (SLAM) is pivotal in robotics, with photorealistic scene reconstruction emerging as a key challenge. To address this, we introduce Computational Alignment for Real-Time Gaussian Splatting SLAM (CaRtGS),…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Dapeng Feng , Zhiqiang Chen , Yizhen Yin , Shipeng Zhong , Yuhua Qi , Hongbo Chen

Text-driven 3D scene editing has attracted considerable interest due to its convenience and user-friendliness. However, methods that rely on implicit 3D representations, such as Neural Radiance Fields (NeRF), while effective in rendering…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Pengcheng Xue , Yan Tian , Qiutao Song , Ziyi Wang , Linyang He , Weiping Ding , Mahmoud Hassaballah , Karen Egiazarian , Wei-Fa Yang , Leszek Rutkowski

Abstract representations of 3D scenes play a crucial role in computer vision, enabling a wide range of applications such as mapping, localization, surface reconstruction, and even advanced tasks like SLAM and rendering. Among these…

计算机视觉与模式识别 · 计算机科学 2024-12-16 Chenggang Yang , Yuang Shi

We propose GaussCtrl, a text-driven method to edit a 3D scene reconstructed by the 3D Gaussian Splatting (3DGS). Our method first renders a collection of images by using the 3DGS and edits them by using a pre-trained 2D diffusion model…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Jing Wu , Jia-Wang Bian , Xinghui Li , Guangrun Wang , Ian Reid , Philip Torr , Victor Adrian Prisacariu

3D Gaussian splatting (3DGS) has recently emerged as an alternative representation that leverages a 3D Gaussian-based representation and introduces an approximated volumetric rendering, achieving very fast rendering speed and promising…

计算机视觉与模式识别 · 计算机科学 2024-08-08 Joo Chan Lee , Daniel Rho , Xiangyu Sun , Jong Hwan Ko , Eunbyung Park

Applying Gaussian Splatting to perception tasks for 3D scene understanding is becoming increasingly popular. Most existing works primarily focus on rendering 2D feature maps from novel viewpoints, which leads to an imprecise 3D language…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Hao Li , Minghan Qin , Zhengyu Zou , Diqi He , Xinhao Ji , Bohan Li , Bingquan Dai , Dingewn Zhang , Junwei Han

3D Gaussian Splatting has shown remarkable capabilities in novel view rendering tasks and exhibits significant potential for multi-view optimization.However, the original 3D Gaussian Splatting lacks color representation for inputs in…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Haoran Wang , Jingwei Huang , Lu Yang , Tianchen Deng , Gaojing Zhang , Mingrui Li

We present GaussExplorer, a framework for embodied exploration and reasoning built on 3D Gaussian Splatting (3DGS). While prior approaches to language-embedded 3DGS have made meaningful progress in aligning simple text queries with Gaussian…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Kim Yu-Ji , Dahye Lee , Kim Jun-Seong , GeonU Kim , Nam Hyeon-Woo , Yongjin Kwon , Yu-Chiang Frank Wang , Jaesung Choe , Tae-Hyun Oh

Visual search in 3D environments requires embodied agents to actively explore their surroundings and acquire task-relevant evidence. However, existing visual search and embodied AI benchmarks, including EQA, typically rely on static…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Koya Sakamoto , Taiki Miyanishi , Daichi Azuma , Shuhei Kurita , Shu Morikuni , Naoya Chiba , Motoaki Kawanabe , Yusuke Iwasawa , Yutaka Matsuo

This study addresses the challenge of generating online 3D Gaussian Splatting (3DGS) models from RGB-only frames. Previous studies have employed dense SLAM techniques to estimate 3D scenes from keyframes for 3DGS model construction.…

计算机视觉与模式识别 · 计算机科学 2025-09-08 Byeonggwon Lee , Junkyu Park , Khang Truong Giang , Soohwan Song

We present GStalker, a 3D audio-driven talking face generation model with Gaussian Splatting for both fast training (40 minutes) and real-time rendering (125 FPS) with a 3$\sim$5 minute video for training material, in comparison with…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Bo Chen , Shoukang Hu , Qi Chen , Chenpeng Du , Ran Yi , Yanmin Qian , Xie Chen

Recent advancements in multi-modal 3D pre-training methods have shown promising efficacy in learning joint representations of text, images, and point clouds. However, adopting point clouds as 3D representation fails to fully capture the…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Haoyuan Li , Yanpeng Zhou , Tao Tang , Jifei Song , Yihan Zeng , Michael Kampffmeyer , Hang Xu , Xiaodan Liang

Conventional geometry-based SLAM systems lack dense 3D reconstruction capabilities since their data association usually relies on feature correspondences. Additionally, learning-based SLAM systems often fall short in terms of real-time…

计算机视觉与模式识别 · 计算机科学 2024-08-22 Zhongche Qu , Zhi Zhang , Cong Liu , Jianhua Yin

Talking head generation is increasingly important in virtual reality (VR), especially for social scenarios involving multi-turn conversation. Existing approaches face notable limitations: mesh-based 3D methods can model dual-person dialogue…

计算机视觉与模式识别 · 计算机科学 2026-01-16 Peng Chen , Xiaobao Wei , Yi Yang , Naiming Yao , Hui Chen , Feng Tian

Reconstructing dynamic driving scenes is essential for developing autonomous systems through sensor-realistic simulation. Although recent methods achieve high-fidelity reconstructions, they either rely on costly human annotations for object…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Carl Lindström , Mahan Rafidashti , Maryam Fatemi , Lars Hammarstrand , Martin R. Oswald , Lennart Svensson

Reconstructing and predicting dynamic 3D scenes from multi-view videos is a foundational task for robotics, AR/VR, and digital twins. Recent physics-informed Gaussian Splatting methods achieve impressive future frame extrapolation but lack…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Denis Gridusov , Maxim Popov , Sergey Kolyubin