English

Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning

Computer Vision and Pattern Recognition 2026-02-02 v3

Abstract

While advanced methods like VACE and Phantom have advanced video generation for specific subjects in diverse scenarios, they struggle with multi-human identity preservation in dynamic interactions, where consistent identities across multiple characters are critical. To address this, we propose Identity-GRPO, a human feedback-driven optimization pipeline for refining multi-human identity-preserving video generation. First, we construct a video reward model trained on a large-scale preference dataset containing human-annotated and synthetic distortion data, with pairwise annotations focused on maintaining human consistency throughout the video. We then employ a GRPO variant tailored for multi-human consistency, which greatly enhances both VACE and Phantom. Through extensive ablation studies, we evaluate the impact of annotation quality and design choices on policy optimization. Experiments show that Identity-GRPO achieves up to 18.9% improvement in human consistency metrics over baseline methods, offering actionable insights for aligning reinforcement learning with personalized video generation.

Keywords

Cite

@article{arxiv.2510.14256,
  title  = {Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning},
  author = {Xiangyu Meng and Zixian Zhang and Zhenghao Zhang and Junchao Liao and Long Qin and Weizhi Wang},
  journal= {arXiv preprint arXiv:2510.14256},
  year   = {2026}
}

Comments

Our project and code are available at https://ali-videoai.github.io/identity_page, https://github.com/alibaba/identity-grpo

R2 v1 2026-07-01T06:40:23.265Z