中文

Proteus-ID:面向视频定制的身份一致且动作连贯

计算机视觉与模式识别 2026-02-04 v2

摘要

视频身份定制旨在给定单个 reference 图像和 text prompt,合成逼真且在时间上连贯的特定主体视频。这一任务 presents 两个 core 挑战:(1) 在保持 identity consistency 的同时与描述的外观和动作对齐,以及 (2) 生成自然流畅的动作而不呈现不切实际的僵硬。为此,我们提出 Proteus-ID,一个 novel diffusion-based 框架,用于 identity-consistent 与 motion-coherent 视频定制。首先,我们提出 Multimodal Identity Fusion (MIF) 模块,通过 Q-Former 将视觉与文本线索统一为 joint identity representation,为扩散模型提供连贯的指导,消除 modality imbalance。其次,我们提出 Time-Aware Identity Injection (TAII) 机制,动态调制 identity conditioning across denoising 步骤,提高 fine-detail 重建。第三,我们提出 Adaptive Motion Learning (AML),一种 self-supervised 策略,基于 optical-flow 提取的 motion heatmaps 对 training loss 进行 reweighting,从而在不额外输入的情况下增强 motion 真实性。为支持该任务,我们构建 Proteus-Bench,包含 200K 组精选 clip 用于 training,以及 150 个来自不同职业与民族的 individuals 用于 evaluation。大量实验表明,Proteus-ID 在 identity preservation、text alignment 与 motion quality 上超越 prior methods,确立了 video identity customization 的新 benchmark。代码与数据已公开于 https://grenoble-zhang.github.io/Proteus-ID/。

关键词

引用

@article{arxiv.2506.23729,
  title  = {Proteus-ID: ID-Consistent and Motion-Coherent Video Customization},
  author = {Guiyu Zhang and Chen Shi and Zijian Jiang and Xunzhi Xiang and Jingjing Qian and Shaoshuai Shi and Li Jiang},
  journal= {arXiv preprint arXiv:2506.23729},
  year   = {2026}
}

备注

SIGGRAPH Asia 2025