English

Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising

Computer Vision and Pattern Recognition 2025-01-07 v1

Abstract

Recent advances in diffusion models have greatly improved text-driven video generation. However, training models for long video generation demands significant computational power and extensive data, leading most video diffusion models to be limited to a small number of frames. Existing training-free methods that attempt to generate long videos using pre-trained short video diffusion models often struggle with issues such as insufficient motion dynamics and degraded video fidelity. In this paper, we present Brick-Diffusion, a novel, training-free approach capable of generating long videos of arbitrary length. Our method introduces a brick-to-wall denoising strategy, where the latent is denoised in segments, with a stride applied in subsequent iterations. This process mimics the construction of a staggered brick wall, where each brick represents a denoised segment, enabling communication between frames and improving overall video quality. Through quantitative and qualitative evaluations, we demonstrate that Brick-Diffusion outperforms existing baseline methods in generating high-fidelity videos.

Keywords

Cite

@article{arxiv.2501.02741,
  title  = {Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising},
  author = {Yunlong Yuan and Yuanfan Guo and Chunwei Wang and Hang Xu and Li Zhang},
  journal= {arXiv preprint arXiv:2501.02741},
  year   = {2025}
}

Comments

ICASSP 2025

R2 v1 2026-06-28T20:57:09.134Z