English

3DFacePolicy: Audio-Driven 3D Facial Animation Based on Action Control

Computer Vision and Pattern Recognition 2025-08-13 v2 Artificial Intelligence Machine Learning Multimedia Sound Audio and Speech Processing

Abstract

Audio-driven 3D facial animation has achieved significant progress in both research and applications. While recent baselines struggle to generate natural and continuous facial movements due to their frame-by-frame vertex generation approach, we propose 3DFacePolicy, a pioneer work that introduces a novel definition of vertex trajectory changes across consecutive frames through the concept of "action". By predicting action sequences for each vertex that encode frame-to-frame movements, we reformulate vertex generation approach into an action-based control paradigm. Specifically, we leverage a robotic control mechanism, diffusion policy, to predict action sequences conditioned on both audio and vertex states. Extensive experiments on VOCASET and BIWI datasets demonstrate that our approach significantly outperforms state-of-the-art methods and is particularly expert in dynamic, expressive and naturally smooth facial animations.

Keywords

Cite

@article{arxiv.2409.10848,
  title  = {3DFacePolicy: Audio-Driven 3D Facial Animation Based on Action Control},
  author = {Xuanmeng Sha and Liyun Zhang and Tomohiro Mashita and Naoya Chiba and Yuki Uranishi},
  journal= {arXiv preprint arXiv:2409.10848},
  year   = {2025}
}
R2 v1 2026-06-28T18:47:09.512Z