English
Related papers

Related papers: GSTalker: Real-time Audio-Driven Talking Face Gene…

200 papers

3D Gaussian splatting-based talking head synthesis has recently gained attention for its ability to render high-fidelity images with real-time inference speed. However, since it is typically trained on only a short video that lacks the…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Junuk Cha , Seongro Yoon , Valeriya Strizhkova , Francois Bremond , Seungryul Baek

Speech-driven talking head generation is a critical yet challenging task with applications in augmented reality and virtual human modeling. While recent approaches using autoregressive and diffusion-based models have achieved notable…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Yihong Lin , Zhaoxin Fan , Xianjia Wu , Lingyu Xiong , Liang Peng , Xiandong Li , Wenxiong Kang , Songju Lei , Huang Xu

Gaussian splatting has emerged as a powerful 3D representation that harnesses the advantages of both explicit (mesh) and implicit (NeRF) 3D representations. In this paper, we seek to leverage Gaussian splatting to generate realistic…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 Ye Yuan , Xueting Li , Yangyi Huang , Shalini De Mello , Koki Nagano , Jan Kautz , Umar Iqbal

We present, GauHuman, a 3D human model with Gaussian Splatting for both fast training (1 ~ 2 minutes) and real-time rendering (up to 189 FPS), compared with existing NeRF-based implicit representation modelling frameworks demanding hours of…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Shoukang Hu , Ziwei Liu

Language-guided 3D scene understanding is important for advancing applications in robotics, AR/VR, and human-computer interaction, enabling models to comprehend and interact with 3D environments through natural language. While 2D…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Anh Thai , Songyou Peng , Kyle Genova , Leonidas Guibas , Thomas Funkhouser

Real-time rendering of high-fidelity and animatable avatars from monocular videos remains a challenging problem in computer vision and graphics. Over the past few years, the Neural Radiance Field (NeRF) has made significant progress in…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Qipeng Yan , Mingyang Sun , Lihua Zhang

Despite recent progress in 3D head avatar generation, balancing identity preservation, i.e., reconstruction, with novel poses and expressions, i.e., animation, remains a challenge. Existing methods struggle to adapt Gaussians to varying…

Graphics · Computer Science 2025-07-25 SeungJun Moon , Hah Min Lew , Seungeun Lee , Ji-Su Kang , Gyeong-Moon Park

In this paper, we introduce GaussianMotion, a novel human rendering model that generates fully animatable scenes aligned with textual descriptions using Gaussian Splatting. Although existing methods achieve reasonable text-to-3D generation…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Gyumin Shim , Sangmin Lee , Jaegul Choo

We propose a novel 3D deepfake generation framework based on 3D Gaussian Splatting that enables realistic, identity-preserving face swapping and reenactment in a fully controllable 3D space. Compared to conventional 2D deepfake approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Wending Liu , Siyun Liang , Huy H. Nguyen , Isao Echizen

Deformable Gaussian Splatting (GS) accomplishes photorealistic dynamic 3-D reconstruction from dense multi-view video (MVV) by learning to deform a canonical GS representation. However, in filmmaking, tight budgets can result in sparse…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Adrian Azzarelli , Nantheera Anantrasirichai , David R Bull

Recent advancements in 3D Gaussian Splatting (3DGS) have unlocked significant potential for modeling 3D head avatars, providing greater flexibility than mesh-based methods and more efficient rendering compared to NeRF-based approaches.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Peizhi Yan , Rabab Ward , Qiang Tang , Shan Du

Constructing a 3D scene capable of accommodating open-ended language queries, is a pivotal pursuit, particularly within the domain of robotics. Such technology facilitates robots in executing object manipulations based on human language…

Neural implicit representations, including Neural Distance Fields and Neural Radiance Fields, have demonstrated significant capabilities for reconstructing surfaces with complicated geometry and topology, and generating novel views of a…

Graphics · Computer Science 2024-02-08 Lin Gao , Jie Yang , Bo-Tao Zhang , Jia-Mu Sun , Yu-Jie Yuan , Hongbo Fu , Yu-Kun Lai

We introduce GenSync, a novel framework for multi-identity lip-synced video synthesis using 3D Gaussian Splatting. Unlike most existing 3D methods that require training a new model for each identity , GenSync learns a unified network that…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Anushka Agarwal , Muhammad Yusuf Hassan , Talha Chafekar

4D content generation has achieved remarkable progress recently. However, existing methods suffer from long optimization times, a lack of motion controllability, and a low quality of details. In this paper, we introduce DreamGaussian4D…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Jiawei Ren , Liang Pan , Jiaxiang Tang , Chi Zhang , Ang Cao , Gang Zeng , Ziwei Liu

We introduce Gaussian Articulated Template Model GART, an explicit, efficient, and expressive representation for non-rigid articulated subject capturing and rendering from monocular videos. GART utilizes a mixture of moving 3D Gaussians to…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Jiahui Lei , Yufu Wang , Georgios Pavlakos , Lingjie Liu , Kostas Daniilidis

Animatable 3D reconstruction has significant applications across various fields, primarily relying on artists' handcraft creation. Recently, some studies have successfully constructed animatable 3D models from monocular videos. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Tingyang Zhang , Qingzhe Gao , Weiyu Li , Libin Liu , Baoquan Chen

3D Gaussian Splatting (3DGS) enables high-fidelity real-time rendering, a key requirement for immersive applications. However, the extension of 3DGS to dynamic scenes remains limitations on the substantial data volume of dense Gaussians and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-01 Jiayu Yang , Weijian Su , Songqian Zhang , Yuqi Han , Jinli Suo , Qiang Zhang

Dynamic reconstruction of deformable tissues in endoscopic video is a key technology for robot-assisted surgery. Recent reconstruction methods based on neural radiance fields (NeRFs) have achieved remarkable results in the reconstruction of…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Weixing Xie , Junfeng Yao , Xianpeng Cao , Qiqin Lin , Zerui Tang , Xiao Dong , Xiaohu Guo

Creating controllable 3D human portraits from casual smartphone videos is highly desirable due to their immense value in AR/VR applications. The recent development of 3D Gaussian Splatting (3DGS) has shown improvements in rendering quality…

Computer Vision and Pattern Recognition · Computer Science 2024-02-07 Alfredo Rivero , ShahRukh Athar , Zhixin Shu , Dimitris Samaras