English

Instruct-Video2Avatar: Video-to-Avatar Generation with Instructions

Computer Vision and Pattern Recognition 2023-06-06 v1

Abstract

We propose a method for synthesizing edited photo-realistic digital avatars with text instructions. Given a short monocular RGB video and text instructions, our method uses an image-conditioned diffusion model to edit one head image and uses the video stylization method to accomplish the editing of other head images. Through iterative training and update (three times or more), our method synthesizes edited photo-realistic animatable 3D neural head avatars with a deformable neural radiance field head synthesis method. In quantitative and qualitative studies on various subjects, our method outperforms state-of-the-art methods.

Keywords

Cite

@article{arxiv.2306.02903,
  title  = {Instruct-Video2Avatar: Video-to-Avatar Generation with Instructions},
  author = {Shaoxu Li},
  journal= {arXiv preprint arXiv:2306.02903},
  year   = {2023}
}

Comments

https://github.com/lsx0101/Instruct-Video2Avatar

R2 v1 2026-06-28T10:56:38.954Z