English

WildActor: Unconstrained Identity-Preserving Video Generation

Computer Vision and Pattern Recognition 2026-03-10 v2

Abstract

Production-ready human video generation requires digital actors to maintain strictly consistent full-body identities across dynamic shots, viewpoints and motions, a setting that remains challenging for existing methods. Prior methods often suffer from face-centric behavior that neglects body-level consistency, or produce copy-paste artifacts where subjects appear rigid due to pose locking. We present Actor-18M, a large-scale human video dataset designed to capture identity consistency under unconstrained viewpoints and environments. Actor-18M comprises 1.6M videos with 18M corresponding human images, covering both arbitrary views and canonical three-view representations. Leveraging Actor-18M, we propose WildActor, a framework for any-view conditioned human video generation. We introduce an Asymmetric Identity-Preserving Attention mechanism coupled with a Viewpoint-Adaptive Monte Carlo Sampling strategy that iteratively re-weights reference conditions by marginal utility for balanced manifold coverage. Evaluated on the proposed Actor-Bench, WildActor consistently preserves body identity under diverse shot compositions, large viewpoint transitions, and substantial motions, surpassing existing methods in these challenging settings.

Cite

@article{arxiv.2603.00586,
  title  = {WildActor: Unconstrained Identity-Preserving Video Generation},
  author = {Qin Guo and Tianyu Yang and Xuanhua He and Fei Shen and Yong Zhang and Zhuoliang Kang and Xiaoming Wei and Dan Xu},
  journal= {arXiv preprint arXiv:2603.00586},
  year   = {2026}
}

Comments

Project Page: https://wildactor.github.io/

R2 v1 2026-07-01T10:57:06.593Z