English

X-Actor: Emotional and Expressive Long-Range Portrait Acting from Audio

Computer Vision and Pattern Recognition 2025-08-06 v1

Abstract

We present X-Actor, a novel audio-driven portrait animation framework that generates lifelike, emotionally expressive talking head videos from a single reference image and an input audio clip. Unlike prior methods that emphasize lip synchronization and short-range visual fidelity in constrained speaking scenarios, X-Actor enables actor-quality, long-form portrait performance capturing nuanced, dynamically evolving emotions that flow coherently with the rhythm and content of speech. Central to our approach is a two-stage decoupled generation pipeline: an audio-conditioned autoregressive diffusion model that predicts expressive yet identity-agnostic facial motion latent tokens within a long temporal context window, followed by a diffusion-based video synthesis module that translates these motions into high-fidelity video animations. By operating in a compact facial motion latent space decoupled from visual and identity cues, our autoregressive diffusion model effectively captures long-range correlations between audio and facial dynamics through a diffusion-forcing training paradigm, enabling infinite-length emotionally-rich motion prediction without error accumulation. Extensive experiments demonstrate that X-Actor produces compelling, cinematic-style performances that go beyond standard talking head animations and achieves state-of-the-art results in long-range, audio-driven emotional portrait acting.

Keywords

Cite

@article{arxiv.2508.02944,
  title  = {X-Actor: Emotional and Expressive Long-Range Portrait Acting from Audio},
  author = {Chenxu Zhang and Zenan Li and Hongyi Xu and You Xie and Xiaochen Zhao and Tianpei Gu and Guoxian Song and Xin Chen and Chao Liang and Jianwen Jiang and Linjie Luo},
  journal= {arXiv preprint arXiv:2508.02944},
  year   = {2025}
}

Comments

Project Page at https://byteaigc.github.io/X-Actor/