English

Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance

Computer Vision and Pattern Recognition 2026-05-18 v1 Artificial Intelligence

Abstract

Video editing poses a significant challenge. While a series of tuning-free methods circumvent the need for extensive data collection and model training, they often underutilize the rich information embedded within noisy latent, leading to unsatisfactory results. To address this, we propose a \textit{tuning-free, instruction-based} video editing framework. We approach video editing from the perspective of noisy latent: we design a Structural Noise Initialization Strategy (SNIS) to secure a superior editing starting point by assigning higher noise levels to edited regions (to facilitate content change) and lower noise levels to unedited regions (to maintain content consistency). We introduce a Noise Guidance Mechanism (NGM), which leverages the video prior in the generative model and effectively integrates rich information within the noisy latent to guide the denoising process, thereby preserving unedited content and overall visual coherence. Experiments show that our proposed method achieves better visual quality and state-of-the-art performance.

Keywords

Cite

@article{arxiv.2605.15533,
  title  = {Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance},
  author = {Song Wu and Xinyu Chen and Qian Wang and Liang Li and Zili Yi and Junlan Feng},
  journal= {arXiv preprint arXiv:2605.15533},
  year   = {2026}
}

Comments

Accepted by ICIP 2026

R2 v1 2026-07-22T07:13:34.894Z