中文

多视图图像提示的多视图扩散用于改进3D生成

计算机视觉与模式识别 2024-04-29 v1

摘要

使用图像作为3D生成的提示相比仅使用文本提示表现出特别强的性能,因为图像为3D生成过程提供了更直观的指导。在这项工作中,我们深入探讨了使用多个图像提示(而非单个图像提示)进行3D生成的潜力。具体来说,我们基于ImageDream(一种新颖的图像提示多视图扩散模型)构建,以支持多视图图像作为输入提示。我们的方法称为MultiImageDream,表明从单图像提示过渡到多图像提示根据各种定量评估指标和定性评估增强了多视图和3D对象生成的性能。这一进步无需微调预训练的ImageDream多视图扩散模型即可实现。

关键词

引用

@article{arxiv.2404.17419,
  title  = {Multi-view Image Prompted Multi-view Diffusion for Improved 3D Generation},
  author = {Seungwook Kim and Yichun Shi and Kejie Li and Minsu Cho and Peng Wang},
  journal= {arXiv preprint arXiv:2404.17419},
  year   = {2024}
}

备注

5 pages including references, 2 figures, 2 tables