English

TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction

Computer Vision and Pattern Recognition 2025-12-03 v3

Abstract

We present TempoMaster, a novel framework that formulates long video generation as next-frame-rate prediction. Specifically, we first generate a low-frame-rate clip that serves as a coarse blueprint of the entire video sequence, and then progressively increase the frame rate to refine visual details and motion continuity. During generation, TempoMaster employs bidirectional attention within each frame-rate level while performing autoregression across frame rates, thus achieving long-range temporal coherence while enabling efficient and parallel synthesis. Extensive experiments demonstrate that TempoMaster establishes a new state-of-the-art in long video generation, excelling in both visual and temporal quality.

Keywords

Cite

@article{arxiv.2511.12578,
  title  = {TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction},
  author = {Yukuo Ma and Cong Liu and Junke Wang and Junqi Liu and Haibin Huang and Zuxuan Wu and Chi Zhang and Xuelong Li},
  journal= {arXiv preprint arXiv:2511.12578},
  year   = {2025}
}

Comments

for more information, see https://scottykma.github.io/tempomaster-gitpage/

R2 v1 2026-07-01T07:39:43.900Z