中文

DipGuava:从单目视频中生成具备 disentangled personalized Gaussian features 的 3D 头部头像

计算机视觉与模式识别 2026-03-31 v1

摘要

虽然最近的 3D 头部头像创建方法试图实现面部动态动画,但往往难以捕获个性化细节,限制了真实感和表达性。为填补这一空白,我们提出了 DipGuava(Disentangled and Personalized Gaussian UV Avatar,即 Disentangled and Personalized Gaussian UV Avatar),这是一种 Novel 的 3D 高斯头像创建方法,成功生成具有个性化属性的头像。DipGuava 是首个显式 disentangle 面部外观的 method,其采用结构化的 two-stage pipeline 显著降低学习歧义并提升重建保真度。在第一阶段,我们学习一个基于几何驱动的 base appearance,以捕获全局面部结构和粗糙的 expression-dependent 变化。在第二阶段,预测第一阶段未捕获的个性化残差细节,包括高频分量和非线性变化特征,如皱纹和细微的皮肤变形。通过 dynamic appearance fusion 将这些 components 融合,将残差细节在变形后整合,确保空间和语义对齐。这种 disentangled 设计使 DipGuava 能够生成具有 photorealistic、identity-preserving 的头像,在视觉质量和 quantitative performance 上均显著优于先前方法,实验结果证明了这一点。

关键词

引用

@article{arxiv.2603.28003,
  title  = {DipGuava: Disentangling Personalized Gaussian Features for 3D Head Avatars from Monocular Video},
  author = {Jeonghaeng Lee and Seok Keun Choi and Zhixuan Li and Weisi Lin and Sanghoon Lee},
  journal= {arXiv preprint arXiv:2603.28003},
  year   = {2026}
}

备注

AAAI 2026