中文
相关论文

相关论文: ChatSplat: 3D Conversational Gaussian Splatting

200 篇论文

The creation of 3D scenes has traditionally been both labor-intensive and costly, requiring designers to meticulously configure 3D assets and environments. Recent advancements in generative AI, including text-to-3D and image-to-3D methods,…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Ziyang Yan , Yihua Shao , Minwen Liao , Siyu Chen , Nan Wang , Muyuan Lin , Jenq-Neng Hwang , Hao Zhao , Fabio Remondino , Lei Li

We propose VideoRFSplat, a direct text-to-3D model leveraging a video generation model to generate realistic 3D Gaussian Splatting (3DGS) for unbounded real-world scenes. To generate diverse camera poses and unbounded spatial extent of…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Hyojun Go , Byeongjun Park , Hyelin Nam , Byung-Hoon Kim , Hyungjin Chung , Changick Kim

Gaussian Splatting (GS) has recently emerged as an efficient representation for rendering 3D scenes from 2D images and has been extended to images, videos, and dynamic 4D content. However, applying style transfer to GS-based…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Kornel Howil , Joanna Waczyńska , Piotr Borycki , Tadeusz Dziarmaga , Marcin Mazur , Przemysław Spurek

We present 3D Spatial MultiModal Memory (M3), a multimodal memory system designed to retain information about medium-sized static scenes through video sources for visual perception. By integrating 3D Gaussian Splatting techniques with…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Xueyan Zou , Yuchen Song , Ri-Zhao Qiu , Xuanbin Peng , Jianglong Ye , Sifei Liu , Xiaolong Wang

This paper presents GaussEdit, a framework for adaptive 3D scene editing guided by text and image prompts. GaussEdit leverages 3D Gaussian Splatting as its backbone for scene representation, enabling convenient Region of Interest selection…

图形学 · 计算机科学 2025-10-01 Zhenyu Shu , Junlong Yu , Kai Chao , Shiqing Xin , Ligang Liu

3D open-vocabulary scene understanding, which accurately perceives complex semantic properties of objects in space, has gained significant attention in recent years. In this paper, we propose GAGS, a framework that distills 2D CLIP features…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Yuning Peng , Haiping Wang , Yuan Liu , Chenglu Wen , Zhen Dong , Bisheng Yang

Jointly estimating camera poses and mapping scenes from RGBD images is a fundamental task in simultaneous localization and mapping (SLAM). State-of-the-art methods employ 3D Gaussians to represent a scene, and render these Gaussians through…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Pengchong Hu , Zhizhong Han

Researchers have conducted many pioneer researches on contactless fingerprints, yet the performance of contactless fingerprint recognition still lags behind contact-based methods primary due to the insufficient contactless fingerprint data…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Yuwei Jia , Yutang Lu , Zhe Cui , Fei Su

Talking head synthesis with arbitrary speech audio is a crucial challenge in the field of digital humans. Recently, methods based on radiance fields have received increasing attention due to their ability to synthesize high-fidelity and…

声音 · 计算机科学 2024-12-12 Yifan Xie , Tao Feng , Xin Zhang , Xiangyang Luo , Zixuan Guo , Weijiang Yu , Heng Chang , Fei Ma , Fei Richard Yu

Realistic scene reconstruction and view synthesis are essential for advancing autonomous driving systems by simulating safety-critical scenarios. 3D Gaussian Splatting excels in real-time rendering and static scene reconstructions but…

计算机视觉与模式识别 · 计算机科学 2024-07-08 Mustafa Khan , Hamidreza Fazlali , Dhruv Sharma , Tongtong Cao , Dongfeng Bai , Yuan Ren , Bingbing Liu

3D Gaussian Splatting has achieved remarkable success in reconstructing both static and dynamic 3D scenes. However, in a scene represented by 3D Gaussian primitives, interactions between objects suffer from inaccurate 3D segmentation,…

图形学 · 计算机科学 2025-06-10 Zeyu Xiao , Zhenyi Wu , Mingyang Sun , Qipeng Yan , Yufan Guo , Zhuoer Liang , Lihua Zhang

We present the first application of 3D Gaussian Splatting in monocular SLAM, the most fundamental but the hardest setup for Visual SLAM. Our method, which runs live at 3fps, utilises Gaussians as the only 3D representation, unifying the…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Hidenobu Matsuki , Riku Murai , Paul H. J. Kelly , Andrew J. Davison

Radiance fields have demonstrated impressive performance in synthesizing lifelike 3D talking heads. However, due to the difficulty in fitting steep appearance changes, the prevailing paradigm that presents facial motions by directly…

计算机视觉与模式识别 · 计算机科学 2024-07-08 Jiahe Li , Jiawei Zhang , Xiao Bai , Jin Zheng , Xin Ning , Jun Zhou , Lin Gu

We introduce ShelfGaussian, an open-vocabulary multi-modal Gaussian-based 3D scene understanding framework supervised by off-the-shelf vision foundation models (VFMs). Gaussian-based methods have demonstrated superior performance and…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Lingjun Zhao , Yandong Luo , James Hays , Lu Gan

3D scene understanding has become an essential area of research with applications in autonomous driving, robotics, and augmented reality. Recently, 3D Gaussian Splatting (3DGS) has emerged as a powerful approach, combining explicit modeling…

计算机视觉与模式识别 · 计算机科学 2025-04-16 Haijie Li , Yanmin Wu , Jiarui Meng , Qiankun Gao , Zhiyao Zhang , Ronggang Wang , Jian Zhang

Language-guided robotic grasping is a rapidly advancing field where robots are instructed using human language to grasp specific objects. However, existing methods often depend on dense camera views and struggle to quickly update scenes,…

机器人学 · 计算机科学 2024-12-04 Junqiu Yu , Xinlin Ren , Yongchong Gu , Haitao Lin , Tianyu Wang , Yi Zhu , Hang Xu , Yu-Gang Jiang , Xiangyang Xue , Yanwei Fu

New era has unlocked exciting possibilities for extending Large Language Models (LLMs) to tackle 3D vision-language tasks. However, most existing 3D multimodal LLMs (MLLMs) rely on compressing holistic 3D scene information or segmenting…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Xiaoyan Wang , Zeju Li , Yifan Xu , Jiaxing Qi , Zhifei Yang , Ruifei Ma , Xiangde Liu , Chao Zhang

Speech-driven facial animation methods usually contain two main classes, 3D and 2D talking face, both of which attract considerable research attention in recent years. However, to the best of our knowledge, the research on 3D talking face…

计算机视觉与模式识别 · 计算机科学 2024-04-22 Yixiang Zhuang , Baoping Cheng , Yao Cheng , Yuntao Jin , Renshuai Liu , Chengyang Li , Xuan Cheng , Jing Liao , Juncong Lin

In this paper, we propose a RGB-D SLAM system that reconstructs a language-aligned dense feature field while sustaining low-latency tracking and mapping. First, we introduce a Top-K Rendering pipeline, a high-throughput and…

机器人学 · 计算机科学 2026-02-10 Seongbo Ha , Sibaek Lee , Kyungsu Kang , Joonyeol Choi , Seungjun Tak , Hyeonwoo Yu

Holistic understanding of urban scenes based on RGB images is a challenging yet important problem. It encompasses understanding both the geometry and appearance to enable novel view synthesis, parsing semantic labels, and tracking moving…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Hongyu Zhou , Jiahao Shao , Lu Xu , Dongfeng Bai , Weichao Qiu , Bingbing Liu , Yue Wang , Andreas Geiger , Yiyi Liao