English

Ingredients: Blending Custom Photos with Video Diffusion Transformers

Computer Vision and Pattern Recognition 2025-03-19 v2

Abstract

This paper presents a powerful framework to customize video creations by incorporating multiple specific identity (ID) photos, with video diffusion Transformers, referred to as Ingredients. Generally, our method consists of three primary modules: (i) a facial extractor that captures versatile and precise facial features for each human ID from both global and local perspectives; (ii) a multi-scale projector that maps face embeddings into the contextual space of image query in video diffusion transformers; (iii) an ID router that dynamically combines and allocates multiple ID embedding to the corresponding space-time regions. Leveraging a meticulously curated text-video dataset and a multi-stage training protocol, Ingredients demonstrates superior performance in turning custom photos into dynamic and personalized video content. Qualitative evaluations highlight the advantages of proposed method, positioning it as a significant advancement toward more effective generative video control tools in Transformer-based architecture, compared to existing methods. The data, code, and model weights are publicly available at: https://github.com/feizc/Ingredients.

Keywords

Cite

@article{arxiv.2501.01790,
  title  = {Ingredients: Blending Custom Photos with Video Diffusion Transformers},
  author = {Zhengcong Fei and Debang Li and Di Qiu and Changqian Yu and Mingyuan Fan},
  journal= {arXiv preprint arXiv:2501.01790},
  year   = {2025}
}
R2 v1 2026-06-28T20:55:26.612Z