English

Projected Representation Conditioning for High-fidelity Novel View Synthesis

Computer Vision and Pattern Recognition 2026-02-13 v1

Abstract

We propose a novel framework for diffusion-based novel view synthesis in which we leverage external representations as conditions, harnessing their geometric and semantic correspondence properties for enhanced geometric consistency in generated novel viewpoints. First, we provide a detailed analysis exploring the correspondence capabilities emergent in the spatial attention of external visual representations. Building from these insights, we propose a representation-guided novel view synthesis through dedicated representation projection modules that inject external representations into the diffusion process, a methodology named ReNoV, short for representation-guided novel view synthesis. Our experiments show that this design yields marked improvements in both reconstruction fidelity and inpainting quality, outperforming prior diffusion-based novel-view methods on standard benchmarks and enabling robust synthesis from sparse, unposed image collections.

Keywords

Cite

@article{arxiv.2602.12003,
  title  = {Projected Representation Conditioning for High-fidelity Novel View Synthesis},
  author = {Min-Seop Kwak and Minkyung Kwon and Jinhyeok Choi and Jiho Park and Seungryong Kim},
  journal= {arXiv preprint arXiv:2602.12003},
  year   = {2026}
}
R2 v1 2026-07-01T10:33:46.383Z