English

A Survey: Spatiotemporal Consistency in Video Generation

Computer Vision and Pattern Recognition 2026-02-19 v2 Artificial Intelligence

Abstract

Video generation aims to produce temporally coherent sequences of visual frames, representing a pivotal advancement in Artificial Intelligence Generated Content (AIGC). Compared to static image generation, video generation poses unique challenges: it demands not only high-quality individual frames but also strong temporal coherence to ensure consistency throughout the spatiotemporal sequence. Although research addressing spatiotemporal consistency in video generation has increased in recent years, systematic reviews focusing on this core issue remain relatively scarce. To fill this gap, this paper views the video generation task as a sequential sampling process from a high-dimensional spatiotemporal distribution, and further discusses spatiotemporal consistency. We provide a systematic review of the latest advancements in the field. The content spans multiple dimensions including generation models, feature representations, generation frameworks, post-processing techniques, training strategies, benchmarks and evaluation metrics, with a particular focus on the mechanisms and effectiveness of various methods in maintaining spatiotemporal consistency. Finally, this paper explores future research directions and potential challenges in this field, aiming to provide valuable insights for advancing video generation technology. The project link is https://github.com/Yin-Z-Y/A-Survey-Spatiotemporal-Consistency-in-Video-Generation.

Keywords

Cite

@article{arxiv.2502.17863,
  title  = {A Survey: Spatiotemporal Consistency in Video Generation},
  author = {Zhiyu Yin and Kehai Chen and Xuefeng Bai and Ruili Jiang and Juntao Li and Hongdong Li and Jin Liu and Yang Xiang and Jun Yu and Min Zhang},
  journal= {arXiv preprint arXiv:2502.17863},
  year   = {2026}
}
R2 v1 2026-06-28T21:56:46.874Z