English

SemanticFace: Semantic Facial Action Estimation via Semantic Distillation in Interpretable Space

Computer Vision and Pattern Recognition 2026-03-19 v2

Abstract

Facial action estimation from a single image is often formulated as predicting or fitting parameters in compact expression spaces, which lack explicit semantic interpretability. However, many practical applications, such as avatar control and human-computer interaction, require interpretable facial actions that correspond to meaningful muscle movements. In this work, we propose SemanticFace, a framework for facial action estimation in the interpretable ARKit blendshape space that reformulates coefficient prediction as structured semantic reasoning. SemanticFace adopts a two-stage semantic distillation paradigm: it first derives structured semantic supervision from ground-truth ARKit coefficients and then distills this knowledge into a multimodal large language model to predict interpretable facial action coefficients from images. Extensive experiments demonstrate that language-aligned semantic supervision improves both coefficient accuracy and perceptual consistency, while enabling strong cross-identity generalization and robustness to large domain shifts, including cartoon faces.

Keywords

Cite

@article{arxiv.2603.14827,
  title  = {SemanticFace: Semantic Facial Action Estimation via Semantic Distillation in Interpretable Space},
  author = {Zejian Kang and Kai Zheng and Yuanchen Fei and Wentao Yang and Hongyuan Zou and Xiangru Huang},
  journal= {arXiv preprint arXiv:2603.14827},
  year   = {2026}
}
R2 v1 2026-07-01T11:21:29.805Z