中文

NavCrafter:从单张图像探索 3D 场景

计算机视觉与模式识别 2026-04-06 v1 人工智能

摘要

从单张图像创建灵活的 3D 场景在 direct 3D 数据获取成本高昂或不实际时至关重要。我们提出了 NavCrafter,一种新框架,通过合成具有 camera controllability 和 temporal-spatial consistency 的 novel-view video sequences 来探索单张图像中的 3D 场景。NavCrafter 利用 video diffusion models 捕获丰富的 3D 先验,采用 geometry-aware expansion 策略逐步扩展 scene coverage。为实现可控的 multi-view synthesis,我们引入了 multi-stage camera control mechanism,通过 dual-branch camera injection and attention modulation 来条件化 diffusion models with diverse trajectories。我们进一步提出了 collision-aware camera trajectory planner 和增强的 3D 高斯雕刻 (3DGS) pipeline,包括 depth-aligned supervision、structural regularization 和 refinement。广泛的实验表明,NavCrafter 在大尺度视角位移下实现了 state-of-the-art 的 novel-view synthesis,显著提高了 3D 重建保真度。

关键词

引用

@article{arxiv.2604.02828,
  title  = {NavCrafter: Exploring 3D Scenes from a Single Image},
  author = {Hongbo Duan and Peiyu Zhuang and Yi Liu and Zhengyang Zhang and Yuxin Zhang and Pengting Luo and Fangming Liu and Xueqian Wang},
  journal= {arXiv preprint arXiv:2604.02828},
  year   = {2026}
}

备注

8 pages accepted by ICRA 2026