中文

基于查询的运动:文本到视频生成中的身份-运动权衡

计算机视觉与模式识别 2025-05-23 v3

摘要

文本到视频扩散模型在根据文本描述生成连贯视频片段方面取得了显著进展。然而,这些模型中运动、结构和身份表征之间的相互作用仍未得到充分探索。在此,我们研究了自注意力查询 (Q) 特征如何同时支配运动、结构和身份,并考察了这些表征相互作用时出现的挑战。我们的分析揭示,Q 不仅影响布局,而且在去噪过程中,Q 对主体身份也有很强的影响,这使得在不产生身份转移副作用的情况下转移运动变得困难。理解这种双重作用使我们能够控制查询特征注入 (Q injection),并展示了两种应用:(1) 一种零样本运动迁移方法——基于 VideoCrafter2 和 WAN 2.1 实现——其效率是现有方法的 10 倍,以及 (2) 一种无需训练的一致性多镜头视频生成技术,其中角色在多个视频镜头中保持身份一致,同时 Q 注入增强了运动保真度。

关键词

引用

@article{arxiv.2412.07750,
  title  = {Motion by Queries: Identity-Motion Trade-offs in Text-to-Video Generation},
  author = {Yuval Atzmon and Rinon Gal and Yoad Tewel and Yoni Kasten and Gal Chechik},
  journal= {arXiv preprint arXiv:2412.07750},
  year   = {2025}
}

备注

(1) Project page: https://research.nvidia.com/labs/par/MotionByQueries/ (2) The methods and results in section 5, "Consistent multi-shot video generation", are based on the arXiv version 1 (v1) of this work. Starting version 2 (v2), we extend and further analyze those findings to efficient motion transfer (3) in v3 we added: results with WAN 2.1, baselines and more quality metrics