Despite the significant progress that has been made in video generative models, existing state-of-the-art methods can only produce videos lasting 5-16 seconds, often labeled "long-form videos". Furthermore, videos exceeding 16 seconds struggle to maintain consistent character appearances and scene layouts throughout the narrative. In particular, multi-subject long videos still fail to preserve character consistency and motion coherence. While some methods can generate videos up to 150 seconds long, they often suffer from frame redundancy and low temporal diversity. Recent work has attempted to produce long-form videos featuring multiple characters, narrative coherence, and high-fidelity detail. We comprehensively studied 32 papers on video generation to identify key architectural components and training strategies that consistently yield these qualities. We also construct a comprehensive novel taxonomy of existing methods and present comparative tables that categorize papers by their architectural designs and performance characteristics.
@article{arxiv.2507.07202,
title = {A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality},
author = {Mohamed Elmoghany and Ryan Rossi and Seunghyun Yoon and Subhojyoti Mukherjee and Eslam Bakr and Puneet Mathur and Gang Wu and Viet Dac Lai and Nedim Lipka and Ruiyi Zhang and Varun Manjunatha and Chien Nguyen and Daksh Dangi and Abel Salinas and Mohammad Taesiri and Hongjie Chen and Xiaolei Huang and Joe Barrow and Nesreen Ahmed and Hoda Eldardiry and Namyong Park and Yu Wang and Jaemin Cho and Anh Totti Nguyen and Zhengzhong Tu and Thien Nguyen and Dinesh Manocha and Mohamed Elhoseiny and Franck Dernoncourt},
journal= {arXiv preprint arXiv:2507.07202},
year = {2025}
}