English

CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Computer Vision and Pattern Recognition 2022-06-01 v1 Computation and Language Machine Learning

Abstract

Large-scale pretrained transformers have created milestones in text (GPT-3) and text-to-image (DALL-E and CogView) generation. Its application to video generation is still facing many challenges: The potential huge computation cost makes the training from scratch unaffordable; The scarcity and weak relevance of text-video datasets hinder the model understanding complex movement semantics. In this work, we present 9B-parameter transformer CogVideo, trained by inheriting a pretrained text-to-image model, CogView2. We also propose multi-frame-rate hierarchical training strategy to better align text and video clips. As (probably) the first open-source large-scale pretrained text-to-video model, CogVideo outperforms all publicly available models at a large margin in machine and human evaluations.

Keywords

Cite

@article{arxiv.2205.15868,
  title  = {CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers},
  author = {Wenyi Hong and Ming Ding and Wendi Zheng and Xinghan Liu and Jie Tang},
  journal= {arXiv preprint arXiv:2205.15868},
  year   = {2022}
}
R2 v1 2026-06-24T11:34:40.030Z