English

AnchorCrafter: Animate Cyber-Anchors Selling Your Products via Human-Object Interacting Video Generation

Computer Vision and Pattern Recognition 2025-06-24 v2

Abstract

The generation of anchor-style product promotion videos presents promising opportunities in e-commerce, advertising, and consumer engagement. Despite advancements in pose-guided human video generation, creating product promotion videos remains challenging. In addressing this challenge, we identify the integration of human-object interactions (HOI) into pose-guided human video generation as a core issue. To this end, we introduce AnchorCrafter, a novel diffusion-based system designed to generate 2D videos featuring a target human and a customized object, achieving high visual fidelity and controllable interactions. Specifically, we propose two key innovations: the HOI-appearance perception, which enhances object appearance recognition from arbitrary multi-view perspectives and disentangles object and human appearance, and the HOI-motion injection, which enables complex human-object interactions by overcoming challenges in object trajectory conditioning and inter-occlusion management. Extensive experiments show that our system improves object appearance preservation by 7.5\% and doubles the object localization accuracy compared to existing state-of-the-art approaches. It also outperforms existing approaches in maintaining human motion consistency and high-quality video generation. Project page including data, code, and Huggingface demo: https://github.com/cangcz/AnchorCrafter.

Keywords

Cite

@article{arxiv.2411.17383,
  title  = {AnchorCrafter: Animate Cyber-Anchors Selling Your Products via Human-Object Interacting Video Generation},
  author = {Ziyi Xu and Ziyao Huang and Juan Cao and Yong Zhang and Xiaodong Cun and Qing Shuai and Yuchen Wang and Linchao Bao and Jintao Li and Fan Tang},
  journal= {arXiv preprint arXiv:2411.17383},
  year   = {2025}
}