English

FA-LAM: Focus-Aware Large Avatar Model for One-Shot 4D Animatable Gaussian Head

Computer Vision and Pattern Recognition 2026-07-23 v1

Abstract

We propose FA-LAM, a Focus-Aware Large Avatar Model for one-shot animatable Gaussian head creation, while simultaneously enabling static 3D and dynamic 4D full-head recovery. The core of our method lies in a thorough analysis of the attention mechanisms and the entangled reconstruction and animation training pipeline adopted by prior state-of-the-art approaches. Our analysis identifies two main factors that compromise the quality of 3D full-head generation: (1) incorrect and noisy attention activations, and (2) conflicts between the tasks of reconstruction and animation. To address the first issue, we introduce a symmetric and semantic attention regularization strategy that leverages the inherent semantics and structural symmetry of human heads. To disentangle the objectives of reconstruction and animation, we develop a novel dual-phase training pipeline that separates the model's capabilities for large-view hallucination and animation into distinct modules. Moreover, we enhance our model to support multi-view and streaming 4D reconstruction in an efficient and memory-friendly manner through a core autoregressive modification with tailored visibility-aware token fusion. Collectively, these innovations enable FA-LAM to reconstruct animatable Gaussian full heads with superior quality, particularly in fine facial regions and large viewing angles.

Cite

@article{arxiv.2607.20922,
  title  = {FA-LAM: Focus-Aware Large Avatar Model for One-Shot 4D Animatable Gaussian Head},
  author = {Yingdong Hu and Yisheng He and Yiming Jiang and Zehong Lin and Steven Hoi and Jun Zhang},
  journal= {arXiv preprint arXiv:2607.20922},
  year   = {2026}
}