English

ViewFusion: Towards Multi-View Consistency via Interpolated Denoising

Computer Vision and Pattern Recognition 2024-03-01 v1

Abstract

Novel-view synthesis through diffusion models has demonstrated remarkable potential for generating diverse and high-quality images. Yet, the independent process of image generation in these prevailing methods leads to challenges in maintaining multiple-view consistency. To address this, we introduce ViewFusion, a novel, training-free algorithm that can be seamlessly integrated into existing pre-trained diffusion models. Our approach adopts an auto-regressive method that implicitly leverages previously generated views as context for the next view generation, ensuring robust multi-view consistency during the novel-view generation process. Through a diffusion process that fuses known-view information via interpolated denoising, our framework successfully extends single-view conditioned models to work in multiple-view conditional settings without any additional fine-tuning. Extensive experimental results demonstrate the effectiveness of ViewFusion in generating consistent and detailed novel views.

Keywords

Cite

@article{arxiv.2402.18842,
  title  = {ViewFusion: Towards Multi-View Consistency via Interpolated Denoising},
  author = {Xianghui Yang and Yan Zuo and Sameera Ramasinghe and Loris Bazzani and Gil Avraham and Anton van den Hengel},
  journal= {arXiv preprint arXiv:2402.18842},
  year   = {2024}
}

Comments

CVPR2024,homepage:https://wi-sc.github.io/ViewFusion.github.io/

R2 v1 2026-06-28T15:04:04.737Z