English

SPEED: One-Step Pixel Diffusion for High-quality Video Frame Interpolation

Multimedia 2026-07-17 v1

Abstract

Despite the success of diffusion models in Video Frame Interpolation (VFI), existing methods still suffer from two critical limitations. First, latent diffusion inevitably loses fine-grained details when reconstructing images from latent representations back to the pixel space. Second, multi-step sampling incurs prohibitive memory consumption and inference latency. To address these issues, we propose SPEED, a one-step pixel diffusion framework for high-quality VFI. Specifically, SPEED employs a progressive multi-stage architecture with dynamic patch scaling to effectively learn multi-scale motion, structural, and appearance representations. Furthermore, we propose a novel Noise-Update-Only Attention mechanism to prevent semantic degradation of the clean condition frames while reducing the computational overhead by nearly 50%. Besides, we introduce a Drift-aware Timestep Sampling strategy coupled with a tailored training objective to directly predict images in the pixel space, enabling one-step inference without compromising the quality of the generated frames. Extensive experiments show that SPEED achieves state-of-the-art performance. On SNU-FILM, SPEED reduces LPIPS by 8.8% while delivering 63.3% faster inference and 10.6% lower memory usage. On challenging 4K benchmarks, it further surpasses prior methods by up to 51.5% in LPIPS.

Cite

@article{arxiv.2607.15585,
  title  = {SPEED: One-Step Pixel Diffusion for High-quality Video Frame Interpolation},
  author = {Zihao Zhang and Haoyu Zhao and Siqian Yang and Yidi Wu and Yudong Jiang and Zuxuan Wu},
  journal= {arXiv preprint arXiv:2607.15585},
  year   = {2026}
}

Comments

ACM MM2026