We present TempoMaster, a novel framework that formulates long video generation as next-frame-rate prediction. Specifically, we first generate a low-frame-rate clip that serves as a coarse blueprint of the entire video sequence, and then progressively increase the frame rate to refine visual details and motion continuity. During generation, TempoMaster employs bidirectional attention within each frame-rate level while performing autoregression across frame rates, thus achieving long-range temporal coherence while enabling efficient and parallel synthesis. Extensive experiments demonstrate that TempoMaster establishes a new state-of-the-art in long video generation, excelling in both visual and temporal quality.
@article{arxiv.2511.12578,
title = {TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction},
author = {Yukuo Ma and Cong Liu and Junke Wang and Junqi Liu and Haibin Huang and Zuxuan Wu and Chi Zhang and Xuelong Li},
journal= {arXiv preprint arXiv:2511.12578},
year = {2025}
}
Comments
for more information, see https://scottykma.github.io/tempomaster-gitpage/