中文

ShapeWords:用于3D形状感知导向文本到图像合成

计算机视觉与模式识别 2025-12-12 v2 人工智能 图形学 机器学习

摘要

我们介绍ShapeWords,这是一种基于3D形状指导和文本提示合成图像的方法。ShapeWords将目标3D形状信息融入专用标记中,这些标记与输入文本一起嵌入,从而有效地将3D形状感知与文本上下文融合,以指导图像合成过程。不同于传统形状指导方法依赖固定视角深度图且常常忽视完整3D结构或文本上下文,ShapeWords生成多样且一致的图像,既反映目标形状的几何,又符合文本描述。实验结果表明,ShapeWords生成的图像在文本合规性、审美合理性方面更优,同时保持3D形状感知。

关键词

引用

@article{arxiv.2412.02912,
  title  = {ShapeWords: Guiding Text-to-Image Synthesis with 3D Shape-Aware Prompts},
  author = {Dmitry Petrov and Pradyumn Goyal and Divyansh Shivashok and Yuanming Tao and Melinos Averkiou and Evangelos Kalogerakis},
  journal= {arXiv preprint arXiv:2412.02912},
  year   = {2025}
}

备注

Project webpage: https://lodurality.github.io/shapewords/ (CVPR 2025 paper), this Author Accepted Manuscript version is made available under the Creative Commons Attribution 4.0 International License (CC BY 4.0) in accordance with the ERC / Horizon Europe open-access mandate