English

Enhance-A-Video: Better Generated Video for Free

Computer Vision and Pattern Recognition 2025-02-28 v3

Abstract

DiT-based video generation has achieved remarkable results, but research into enhancing existing models remains relatively unexplored. In this work, we introduce a training-free approach to enhance the coherence and quality of DiT-based generated videos, named Enhance-A-Video. The core idea is enhancing the cross-frame correlations based on non-diagonal temporal attention distributions. Thanks to its simple design, our approach can be easily applied to most DiT-based video generation frameworks without any retraining or fine-tuning. Across various DiT-based video generation models, our approach demonstrates promising improvements in both temporal consistency and visual quality. We hope this research can inspire future explorations in video generation enhancement.

Keywords

Cite

@article{arxiv.2502.07508,
  title  = {Enhance-A-Video: Better Generated Video for Free},
  author = {Yang Luo and Xuanlei Zhao and Mengzhao Chen and Kaipeng Zhang and Wenqi Shao and Kai Wang and Zhangyang Wang and Yang You},
  journal= {arXiv preprint arXiv:2502.07508},
  year   = {2025}
}
R2 v1 2026-06-28T21:40:10.768Z