English

DEGSTalk: Decomposed Per-Embedding Gaussian Fields for Hair-Preserving Talking Face Synthesis

Computer Vision and Pattern Recognition 2024-12-31 v1 Human-Computer Interaction

Abstract

Accurately synthesizing talking face videos and capturing fine facial features for individuals with long hair presents a significant challenge. To tackle these challenges in existing methods, we propose a decomposed per-embedding Gaussian fields (DEGSTalk), a 3D Gaussian Splatting (3DGS)-based talking face synthesis method for generating realistic talking faces with long hairs. Our DEGSTalk employs Deformable Pre-Embedding Gaussian Fields, which dynamically adjust pre-embedding Gaussian primitives using implicit expression coefficients. This enables precise capture of dynamic facial regions and subtle expressions. Additionally, we propose a Dynamic Hair-Preserving Portrait Rendering technique to enhance the realism of long hair motions in the synthesized videos. Results show that DEGSTalk achieves improved realism and synthesis quality compared to existing approaches, particularly in handling complex facial dynamics and hair preservation. Our code will be publicly available at https://github.com/CVI-SZU/DEGSTalk.

Keywords

Cite

@article{arxiv.2412.20148,
  title  = {DEGSTalk: Decomposed Per-Embedding Gaussian Fields for Hair-Preserving Talking Face Synthesis},
  author = {Kaijun Deng and Dezhi Zheng and Jindong Xie and Jinbao Wang and Weicheng Xie and Linlin Shen and Siyang Song},
  journal= {arXiv preprint arXiv:2412.20148},
  year   = {2024}
}

Comments

Accepted by ICASSP 2025

R2 v1 2026-06-28T20:50:38.740Z